遇见数据集

Inferencelab/RUEmoCorp

收藏
Hugging Face2026-05-18 更新2026-05-31 收录
官方服务:

资源简介:

RUEmoCorp(罗马乌尔都语情感语料库)是一个大规模、人工标注、专家注释的罗马乌尔都语社交媒体和对话文本数据集,标注了7种情感类别:快乐、愤怒、悲伤、恐惧、厌恶、惊讶和无情感。数据收集自巴基斯坦社交媒体平台和WhatsApp对话,经过严格的多阶段标注流程,由四名来自巴基斯坦独立大学的专家标注者进行标注。数据集通过交互标注者一致性(IAA)验证,在700个样本的基准测试中,Fleiss κ = 0.6588,平均配对Cohens κ = 0.6597,表明标注一致性较高。该数据集填补了罗马乌尔都语领域缺乏大规模、公开可访问、IAA验证情感语料库的空白,适用于情感分类、低资源NLP和情感计算任务。数据集包含约28,000个样本,情感类别分布大致平衡,永久存档于哈佛Dataverse,采用CC BY 4.0许可。

RUEmoCorp (Roman Urdu Emotional Corpus) is a large-scale, manually annotated and expert-curated textual dataset of social media and conversational content in Roman Urdu, with seven annotated emotion categories: joy, anger, sadness, fear, disgust, surprise, and neutral. The data was collected from Pakistani social media platforms and WhatsApp conversations, and underwent a strict multi-stage annotation process conducted by four expert annotators from the Independent University of Pakistan. The dataset was validated using Inter-Annotator Agreement (IAA) metrics; in a benchmark test of 700 samples, the Fleiss’ κ score was 0.6588 and the average pairwise Cohen’s κ score was 0.6597, indicating high annotation consistency. This dataset fills the gap in the Roman Urdu domain where large-scale, publicly accessible, IAA-validated emotional corpora are scarce, and is applicable to tasks including sentiment classification, low-resource NLP, and affective computing. The dataset contains approximately 28,000 samples with roughly balanced distribution across emotion categories, and is permanently archived on Harvard Dataverse under the CC BY 4.0 license.

提供机构:
Inferencelab
二维码
社区交流群
二维码
科研交流群
商业服务