Khubaib01/RUEmoCorp
收藏资源简介:
RUEmoCorp(罗马乌尔都语情感语料库)是一个大规模、手动整理、专家标注的数据集,包含罗马乌尔都语的社交媒体和对话文本,标注了7种情感类别:喜悦、愤怒、悲伤、恐惧、厌恶、惊讶和无情感。它是roman-urdu-emotion-xlmr-v2模型的训练语料,该模型是罗马乌尔都语最高准确率的开源情感分类器,达到Macro F1 = 0.9896。数据收集自巴基斯坦社交媒体平台和WhatsApp对话,并经过由来自三所独立巴基斯坦大学的四位专家标注者进行的严格多阶段标注过程。在一个700样本的基准测试中,标注者间一致性(IAA)研究显示Fleiss κ = 0.6588和Mean Pairwise Cohens κ = 0.6597,表明具有实质性一致性——这对于低资源、拼写不规律语言中的7类情感标注任务是一个强有力的结果。情感分类采用Ekman的六种基本情感,并增加了一个无情感类别用于情感中性的话语——这是先前罗马乌尔都语情感工作中缺失的刻意设计选择,之前仅使用四或六种类别。省略中性类别会迫使分类器为中性文本分配情感标签,从而增加部署系统中的误报率。该数据集填补了一个已记录的空白:在此发布之前,没有大规模、公开可访问、IAA验证的情感语料库适用于罗马乌尔都语,尽管罗马乌尔都语是全球超过2.3亿乌尔都语使用者的主要数字书写模式。RUEmoCorp永久存档于哈佛Dataverse(doi:10.7910/DVN/BPWHOZ)并在CC BY 4.0许可下发布。
RUEmoCorp (Roman-Urdu Emotion Corpus) is a large-scale, manually curated, expert-annotated dataset containing social media and conversational texts in Roman Urdu, annotated with seven emotion categories: joy, anger, sadness, fear, disgust, surprise, and neutral emotion. It serves as the training corpus for the roman-urdu-emotion-xlmr-v2 model, which is the open-source emotion classifier with the highest accuracy for Roman Urdu, achieving a Macro F1 score of 0.9896. The data was collected from Pakistani social media platforms and WhatsApp conversations, and underwent a rigorous multi-stage annotation process conducted by four expert annotators from three independent Pakistani universities. In a benchmark test of 700 samples, an inter-annotator agreement (IAA) study yielded a Fleiss’ κ of 0.6588 and a mean pairwise Cohen’s κ of 0.6597, indicating substantial agreement — a strong result for a seven-category emotion annotation task in low-resource, orthographically irregular languages. The emotion classification adopts Ekman’s six basic emotions, with an added neutral emotion category for affectively neutral utterances — a deliberate design choice missing from prior Roman Urdu emotion work, which previously only employed four or six categories. Omitting the neutral category forces classifiers to assign emotion labels to neutral texts, which increases the false positive rate in deployed systems. This dataset fills a well-documented gap: prior to its release, there were no large-scale, publicly accessible, IAA-validated emotion corpora available for Roman Urdu, despite Roman Urdu being the primary digital writing system for over 230 million Urdu speakers globally. RUEmoCorp is permanently archived in Harvard Dataverse (doi:10.7910/DVN/BPWHOZ) and released under the CC BY 4.0 license.



