Inferencelab/RomanUrdu-NLP-Sentiment-Corpus
收藏资源简介:
RomanUrdu-NLP-Sentiment-Corpus是一个最大的公开罗马乌尔都语情感分析数据集,包含134,052个标记文本样本,收集自聊天和社交媒体平台。该数据集设计具有对俚语、缩写和拼写变体的鲁棒性,通过LLM辅助标记和人工验证确保高质量,情感类别(积极、中性、消极)平衡,适用于研究和实际NLP应用。它支持情感分析、低资源语言NLP、代码混合和俚语感知文本建模以及社交媒体意见挖掘等领域的研究。数据集结构包括两列:message(罗马乌尔都语文本)和label(情感类别)。统计数据包括总样本数、唯一消息数、类别分布(积极28%、中性32%、消极40%)以及消息长度统计(平均13.55词、66.62字符)。标注方法结合了LLM辅助和人工验证,确保可扩展性和一致性。数据集还包括本地俚语、非正式表达和混合英语,使其适用于真实世界部署、聊天机器人、社交媒体分析和低资源语言研究。
RomanUrdu-NLP-Sentiment-Corpus is the largest publicly available Roman Urdu sentiment analysis dataset, containing 134,052 labeled text samples collected from chats and social media platforms. The dataset is designed to be robust to slang, abbreviations, and spelling variations, with high-quality annotation through LLM-assisted labeling and human validation, balanced across sentiment classes (Positive, Neutral, Negative), and suitable for research and real-world NLP applications. It supports research in sentiment analysis, low-resource language NLP, code-mixed and slang-aware text modeling, and social media opinion mining. The dataset structure includes two columns: message (Roman Urdu text) and label (sentiment class). Statistics cover total samples, unique messages, class distribution (Positive 28%, Neutral 32%, Negative 40%), and message length statistics (mean 13.55 words, 66.62 characters). The annotation methodology combines LLM assistance and human review for scalability and consistency. The dataset incorporates local slang, informal expressions, and mixed English, making it suitable for real-world deployment, chatbots, social media analysis, and low-resource language research.



