RUEmoCorp
收藏资源简介:
RUEmoCorp: A Large-Scale Roman Urdu Emotion Corpus RUEmoCorp is a large-scale emotion classification corpus for Roman Urdu, the informal phonetically transliterated writing style widely used across Pakistani social media, messaging applications, and online communities. Roman Urdu remains severely underrepresented in natural language processing research despite being one of the dominant written forms of Urdu in digital communication. Unlike standard Urdu written in Nastaliq script, Roman Urdu has no standardized orthography, exhibits substantial spelling variation, and frequently contains code-mixed English expressions, making emotion recognition particularly challenging. To address this gap, RUEmoCorp provides both a formally annotated benchmark dataset and a large-scale raw corpus designed to support research in low-resource multilingual NLP, affective computing, cross-lingual transfer learning, and Roman Urdu language understanding. Dataset Components Training Corpus: A curated label-balanced subset of approximately 28,000 annotated samples used to train the companion transformer-based emotion classification model released alongside this dataset. - The 162k dataset, and ground truth will be released separately. Emotion Taxonomy RUEmoCorp adopts Paul Ekman’s six basic emotion categories with an additional neutral category for emotionally ambiguous or non-affective utterances: joy – happiness, excitement, delight anger – frustration, hostility, rage sadness – grief, disappointment, sorrow fear – anxiety, uncertainty, dread disgust – contempt, revulsion, strong dislike surprise – astonishment, unexpected reactions none – emotionally neutral or ambiguous utterances Data Sources The corpus was collected from naturally occurring Roman Urdu communication contexts, including: Public Pakistani social media posts, comments, and discussion threads Anonymized WhatsApp group conversations contributed by consenting participants All personally identifiable information including names, phone numbers, and URLs was removed or anonymized prior to inclusion in the dataset. Annotation Methodology The benchmark subset was independently labeled by four annotators from: Bahauddin Zakariya University (BZU), Multan COMSATS University Islamabad (CUI) Emerson University Multan (EUM) Annotators were native Urdu speakers and active users of Roman Urdu in digital communication. Annotation followed a structured protocol including: Detailed annotation guidelines with Roman Urdu examples Ground-truth-guided calibration sessions Independent single-label emotion annotation Confidence scoring and secondary-label recording Majority-vote conflict resolution Inter-annotator agreement was formally evaluated using Fleiss’ Kappa and pairwise Cohen’s Kappa: Fleiss’ Kappa: κ = 0.6588 (Substantial Agreement) Mean Pairwise Cohen’s Kappa: κ = 0.6597 Total Annotated Samples: 700 Full Agreement (4/4): 49.7% Majority Agreement (3/4): 34.4% Ambiguous Samples (2–2 split): 15.9% The observed agreement levels are considered strong for subjective affective annotation tasks and are comparable to established multilingual emotion datasets. Companion Model A transformer-based companion model, khubaib01/roman-urdu-emotion-xlmr-v2, is released alongside RUEmoCorp. The model extends XLM-RoBERTa with a custom two-layer MLP classification head for seven-class emotion classification. The model achieves: Macro F1: 0.9896 Weighted F1: 0.9896 Accuracy: 0.9896 Experimental comparisons against mBERT, TF-IDF + SVM, Logistic Regression, and FastText baselines demonstrate consistent improvements from the proposed architecture. Intended Use RUEmoCorp is intended to support research in: Emotion classification for Roman Urdu Low-resource multilingual NLP Cross-lingual transfer learning Affective computing South Asian social media analysis Code-mixed language understanding Ethical Considerations The dataset was curated with anonymization procedures to remove personally identifiable information. Researchers should not use this dataset for surveillance, profiling, or monitoring of individuals based on inferred emotional states. Users should also consider the following limitations: Predominantly Pakistani sociolinguistic context Natural class imbalance in the raw corpus Subjective ambiguity inherent in emotion annotation Temporal evolution of online language usage patterns License RUEmoCorp is released under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Users are free to share and adapt the material for any purpose provided appropriate attribution is given. Citation If you use RUEmoCorp in your research, please cite: Ahmad, M. K., & Faisal, K. (2025). RUEmoCorp: Roman Urdu Emotion Corpus [Data set]. Harvard Dataverse. Contributors Muhammad Khubaib Ahmad – Core Researcher, Lead Engineer, Project Administration, Model Development Khadija Faisal – Data Manager, Annotation Coordination, Annotator Muzammil Shadab – Annotator...



