Dataset for "A BERT-based Multi-task Learning Framework for Social Media Opinion Mining with Risk Detection"
收藏资源简介:
This dataset accompanies the paper "A BERT-based Multi-task Learning Framework for Social Media Opinion Mining with Risk Detection". It contains 4,239 Chinese campus public opinion text samples, generated using a template-based approach for multi-task text classification research combining topic classification and sentiment analysis. The dataset contains 8 topic categories (Complaint, Academic, Social, Entertainment, Help Seeking, Hobbies, Campus Events, Resource Trading) and 3 sentiment categories (Negative, Positive, Neutral). Due to the template-based generation approach, the dataset contains repeated sentence patterns, which is intentional to simulate the repetitive nature of real-world campus opinion text. The dataset contains approximately 1,323 unique sentence patterns, with the remaining samples generated from the same templates with minor variations. Samples were generated in 7 batches; batches 1-6 share similar sentence patterns for the same topic-sentiment combinations, while batch 7 focuses on positive sentiment samples. To verify the validity of the template-generated labels, two annotators independently labeled 89 overlapping samples without seeing the template labels. Inter-annotator agreement was almost perfect (Cohen's Kappa = 0.916 for sentiment, 1.000 for topic), confirming that the template-generated labels are consistent with human judgment. Files: - 1.csv ~ 7.csv: raw batch files (generated in 7 batches) - merged.csv: 4,239 samples with columns (label, review, theme) - annotator_A.csv: 89 samples annotated by Annotator A - annotator_B.csv: 89 samples annotated by Annotator B License: CC BY 4.0



