Hybrid Human-AI Annotated Thoughts Dataset
收藏资源简介:
Title A Novel Human-AI Hybrid Annotated Dataset of Young Adult Thoughts: Collection, Annotation, and Active Learning-Based Labeling Process 1. Data Collection A total of 5,125 thoughts were initially collected to populate the dataset: 3,587 thoughts were sourced from Google Form responses by students and staff of G.L. Bajaj Institute of Technology & Management and Sharda University, Greater Noida, UP, India. 1,538 thoughts were generated using GPT-4 (ChatGPT), employing a controlled prompt-engineering strategy to cover all four cognitive categories (Positive, Negative, Necessary, Peripheral) aligned with the research objectives. During initial quality screening, 107 responses from the Google Form (human-sourced) data were rejected due to irrelevance, duplication, or incompleteness.Final working dataset for annotation: 5,018 thoughts (3,480 human-collected + 1,538 AI-generated). 2. Primary Annotation Phase From the 5,018 thoughts, a sample of 2,050 thoughts (1,435 human-sourced, 615 AI-generated) was selected for in-depth human annotation: These thoughts were distributed to two independent annotation groups, each following strict guidelines (see attached “Labeling of Unlabeled Thoughts By Human Annotators” diagram). For human data: if ≥2 of 3 annotators agreed on thought category, the label was retained. Otherwise, the thought was discarded. 74 thoughts (from the 1,435 human subset) were rejected due to annotation disagreement. All 615 AI-generated thoughts were successfully verified and labeled by the annotators. Result: 1,976 high-quality, labeled thoughts for direct supervised machine learning use. 3. Formation of Labeled and Unlabeled Pools After annotation: Labeled set: 1,976 fully labeled thoughts (1,361 human, 615 AI-verified). Unlabeled pool: 2,968 thoughts remained (2,045 human-sourced, 923 AI-generated), pending further labeling. 4. Active Learning and Machine Annotation To maximize annotation efficiency and dataset completeness, an active learning workflow was implemented, as visualized in the attached process diagram and results table: The labeled dataset (1,976 thoughts) was split into 80:20 ratio for ML model training (SVM, Logistic Regression, Naive Bayes, Random Forest) and testing. Support Vector Machine (SVM) achieved the best baseline accuracy (89.6%) and was selected for active learning. Unlabeled thoughts were divided into four working sets (744, 744, 740, 740). For each iteration: The updated SVM variant predicted labels and assigned a confidence score to each unlabeled example. Thoughts with <40% confidence were manually sent to the annotator groups for verification. Validated thoughts ("assets") were added to the training set to further improve SVM accuracy. This iterative process produced six SVM generations (SVM_1...SVM_6) and incrementally increased labeled data after each cycle. Remaining unlabeled thoughts after these cycles (2,491) were finally annotated using the last and most robust model (SVM_6). All annotation rounds, human or model-driven, are detailed in the attached “Active Learning and Results” table for transparent tracking. 5. Final Dataset Composition Total final labeled thoughts: 4,944 Human-annotated: 2,453 (direct manual verification during initial and active learning cycles) Model-predicted: 2,491 (SVM_6 auto-annotation, validated in earlier cycles) All thoughts (both human and AI-original) passed at least one round of human or model scrutiny before inclusion. 6. Documentation, Reproducibility, and Data Structure All data processing, annotation criteria, model versioning, and prompt engineering details are documented for full reproducibility. Attached diagrams ("Labeling of Unlabeled Thoughts By Human Annotators", "Active Learning Implementation and Results Table") illustrate core decision flows and dataset evolution. Metadata for each thought includes: unique ID,final thought text, and category label. 7. Ethical Statement Participation was voluntary and anonymized. No sensitive personal information was collected.AI-generated content was produced under controlled prompts and fully validated by human experts to avoid synthetic data bias or distortion.



