遇见数据集

Consensus-labeled Emotion Dataset for Korean Full-Length Novels

收藏
Zenodo2026-07-27 更新2026-08-01 收录
官方服务:

资源简介:

# Consensus-labeled Emotion Dataset for Korean Full-Length Novels ## 1. OverviewThis dataset contains the consensus-labeled emotion data for two classic Korean full-length novels: *Sohyeonseongnok* (소현성록) and its sequel *Sossi Samdaerok* (소씨삼대록). The dataset was constructed to mitigate the subjective bias of individual annotators and the contextual limitations of AI predictions. It provides sentence-level emotion labels derived through a "Human-in-the-loop" pipeline, combining quantitative predictions from deep learning models and qualitative annotations from experts in classic Korean literature using a majority-rule principle. ## 2. Dataset Files* `Sohyeonseongnok_consensus-labeled_data.csv`: Contains 5,282 sentences from *Sohyeonseongnok* with individual annotations and the final consensus label.* `Sossi-samdaerok_consensus-labeled_data.csv`: Contains 15,426 sentences from *Sossi-Samdaerok* with the final consensus emotion labels based on the predictions of 8 models and expert verification. ## 3. Data Dictionary (Column Descriptions)### 3.1. Sohyeonseongnok Dataset (`Sohyeonseongnok_consensus-labeled_data.csv`)This dataset contains the text and emotion labels annotated by one AI model and two human experts, along with the final consensus.* `ID`: The sequential identifier for each sentence.* `Text`: The original sentence from the novel (in Korean).* `v.1.1.1`: The emotion label predicted by the initial AI model trained on the *Changseongamuirok* dataset.* `Yoon`: The qualitative emotion label annotated by Expert 1 (Yoon).* `Kwon`: The qualitative emotion label annotated by Expert 2 (Kwon).* `Label comparison`: The intermediate result showing the agreement status among the AI prediction and two expert annotations.* `consensus-label`: The final emotion label determined by the majority rule. If all three sources disagreed, the final label was assigned by the researcher through contextual verification. ### 3.2. Sossi Samdaerok Dataset (`Soccisamdaerok_consensus-labeled_data.csv`)This dataset contains emotion predictions from 8 distinct AI models (differentiated by their training datasets) and the final consensus label.* `ID`: The sequential identifier for each sentence.* `Text`: The original sentence from the novel (in Korean).* `v.1.1.1`: Predicted label by the model trained on *Changseongamuirok* qualitative data.* `v.1.1.2`: Predicted label by the model trained on Expert 2 (Kwon)'s qualitative data.* `v.1.1.3`: Predicted label by the model trained on Expert 1 (Yoon)'s qualitative data.* `v.1.1.4`: Predicted label by the model trained on both experts' (Kwon + Yoon) data.* `v.1.1.5`: Predicted label by the model trained on *Changseongamuirok* + Expert 2 (Kwon) data.* `v.1.1.6`: Predicted label by the model trained on *Changseongamuirok* + Expert 1 (Yoon) data.* `v.1.1.7`: Predicted label by the model trained on *Changseongamuirok* + both experts' data.* `v.1.1.8`: Predicted label by the model trained on the newly established *Sohyeonseongnok* consensus-labeled dataset.* `Label comparison`: The intermediate result tracking the agreement status (e.g., number of matching predictions) among the 8 AI models.* `consensus-label`: The final comprehensive emotion label. Determined by the majority rule (agreement of 5 or more models), with human expert verification for discrepancies depending on narrative contexts. ## 4. Emotion Classification System (18 Categories)The dataset utilizes an 18-category emotion classification system optimized for classic Korean full-length novels. This system reclassifies the original 44 KOTE emotion categories into 18 broader labels based on narrative contexts: 1. **없음 (None)**: Represents the absence of specific emotions.2. **불만 (Dissatisfaction)**: Encompasses dissatisfaction, irritation, distrust, feeling fed up, feeling pathetic, feeling preposterous, and resentment.3. **감동/감사 (Admiration/Gratitude)**: Encompasses admiration, gratitude, and respect.4. **신뢰/호의 (Trust/Favour)**: Encompasses trust, favor, welcome, reassurance, and feeling fortunate/relieved.5. **슬픔 (Sadness)**: Encompasses sadness, sorrow, and exhaustion.6. **분노/혐오 (Anger/Disgust)**: Encompasses anger, disgust, contempt, and shock.7. **기대 (Expectation)**: Encompasses expectancy, interest, and hope/desire.8. **애정 (Affection)**: Encompasses care and affection, and love.9. **기쁨 (Joy)**: Encompasses joy, happiness, comfort, pride, and excitement.10. **우쭐댐/무시함 (Arrogance)**: Encompasses arrogance, disregard, mockery, and teasing/harassment.11. **안타까움/실망 (Disappointment)**: Encompasses disappointment.12. **불쌍함/연민 (Compassion)**: Encompasses compassion and pity.13. **수치/죄책 (Shame/Guilt)**: Encompasses shame, guilt, despair, and defeat/self-hatred.14. **두려움 (Fear)**: Encompasses fear, anxiety, worry/concern, and apprehension.15. **깨달음 (Realization)**: Encompasses realization and lesson/admonition16. **당황/난처 (Embarrassment/Reluctance)**: Encompasses embarrassment, reluctance, and finding something strange/baffling.17. **비장함 (Resolute)**: Encompasses resolute determination.18. **놀람 (Surprise)**: Encompasses surprise. ## 5. MethodologyThe dataset employs two distinct pipelines for consensus depending on the source material: **For *Sohyeonseongnok*:**1. **Independent Annotation(AI and Experts):** Emotions in *Sohyeonseongnok* were predicted using a deep learning model (v.1.1.1) trained on emotion data of *Changseongamuirok* identified by researcher, while two experts independently annotated the emotions of *Sohyeonseongnok*.2. **Majority Rule Consensus:** Labels were aggregated among the 3 sources. If at least two sources agreed, the label was maintained. 3. **Researcher Verification:** For sentences where all 3 sources provided completely different labels, a researcher manually reviewed the narrative context to assign the final `consensus-label`. **For *Sossi Samdaerok*:**1. **Construction of 8 Models with Diverse Training Conditions:** Sentences were predicted by 8 distinct deep learning emotion prediction models trained on various combinations of qualitative datasets.2. **Multi-Model Majority Rule Consensus:** Labels were aggregated among the 8 models. If 5 or more models agreed, the label was maintained.3. **Final Verification Based on Narrative Context:** For sentences that failed to reach the 5-model consensus threshold, or required contextual reassessment despite reaching the threshold, a researcher manually reviewed and assigned the final `consensus-labeled data`. For any questions regarding this dataset, please contact the primary author.

提供机构:
Zenodo
创建时间:
2026-07-27
二维码
社区交流群
二维码
科研交流群
商业服务