UrduABSADataset
收藏资源简介:
UrduABSADataset是一个用于方面级情感分析(ABSA)的乌尔都语推特数据集,包含约30万条标注推文。数据来源于Twitter(X)平台,覆盖了政治、体育、娱乐、新闻、COVID-19、事件、教育、医疗保健等16个多样化主题。该数据集的一个关键特点是同时提供了原始版本和清洗版本:原始版本完整保留了社交媒体文本中特有的语言特征,如缩写、非正式表达和特定写作风格,以支持真实世界应用的鲁棒性研究;清洗版本则经过标准化处理,并提供了五个层面的精细标注,包括方面词、观点词、情感极性(针对方面)、方面类别以及类别情感极性。该数据集的创建旨在解决乌尔都语作为低资源语言在细粒度情感分析任务上标注数据匮乏的问题。它适用于多种自然语言处理任务,包括方面级情感分析、文本分类、观点挖掘、主题建模和信息检索,可作为微调乌尔都语预训练模型(如UrduBERT、mT5、XLM-R)或评估相关工具性能的基准资源。数据集的标注采用了人工与大语言模型(LLM)相结合的混合方法。
UrduABSADataset is an Urdu Twitter dataset for aspect-based sentiment analysis (ABSA), containing approximately 300,000 annotated tweets. The data is sourced from the Twitter (X) platform, covering 16 diverse topics such as politics, sports, entertainment, news, COVID-19, events, education, and healthcare. A key feature of this dataset is that it provides both raw and cleaned versions: the raw version fully retains language-specific characteristics of social media text, such as abbreviations, informal expressions, and specific writing styles, to support robustness research in real-world applications; the cleaned version is standardized and offers fine-grained annotations at five levels, including aspect terms, opinion terms, sentiment polarity (for aspects), aspect categories, and category sentiment polarity. The creation of this dataset aims to address the scarcity of annotated data for fine-grained sentiment analysis tasks in Urdu as a low-resource language. It is suitable for various natural language processing tasks, including aspect-based sentiment analysis, text classification, opinion mining, topic modeling, and information retrieval, and can serve as a benchmark resource for fine-tuning Urdu pre-trained models (such as UrduBERT, mT5, XLM-R) or evaluating the performance of related tools. The annotation of the dataset employs a hybrid method combining human effort and large language models (LLMs).
数据集概述:Urdu Tweet ABSA Dataset (UrduABSADataset)
- 数据集名称: Urdu Tweet ABSA Dataset (UrduABSADataset)
- 数据集规模: 约 300,000 条乌尔都语推文(位于 100K < n < 1M 范围)
- 语言: 乌尔都语 (ur),使用乌尔都文字
- 许可证: Creative Commons Attribution 4.0 International (CC BY 4.0)
- 策划者: Zoya, Seemab Latif
- 发布者: Zoya
数据集来源
- 数据来源: 通过 Twitter API 收集,初始约 5000 万条乌尔都语推文,经预处理和清洗后保留约 30 万条。
- 主题覆盖: 涵盖 16 个多样化主题,包括政治、体育、娱乐、新闻、COVID-19、事件、教育与医疗保健。
- 相关论文:
- 基于大语言模型的乌尔都语方面级情感分析伪标签方法(Data Intelligence, 2025)
- 乌尔都语方面级情感分析:使用大语言模型的混合数据集增强(IEEE HONET, 2025)
- 基于自动标签的乌尔都语推文 LDA 和 NMF 主题模型分析(IEEE Access, 2021)
- 基于统计和异常检测方法评估乌尔都语处理工具(ACM TALLIP, 2023)
数据集用途
- 直接用途:
- 方面级情感分析 (ABSA)
- 细粒度意见挖掘
- 低资源语言文本分类
- 主题建模与信息检索
- 乌尔都语处理工具评估
- 可用于微调 Transformer 模型(如 UrduBERT、mT5、XLM-R)或作为 ABSA 基准
- 超出范围用途: 文本分析以外的任务(如图像、语音或视频处理)
数据集结构
数据集提供两个版本:
- 原始版本: 保留社交媒体特有的语言特征,如缩写、非正式表达。
- 清洗与标准化版本: 提供五个层级的标注:方面词、意见词、情感极性、方面类别、方面类别极性。
标注信息
- 标注类型: 方面词、意见词、情感极性、方面类别、方面类别极性
- 标注方式: 人工与基于大语言模型的混合标注
创建动机
乌尔都语作为低资源语言,缺乏可用于细粒度情感分析的公开标注数据集。现有资源多为文档级或句子级情感分析,缺少方面和意见标注。该数据集旨在填补这一空白,支持乌尔都语社交媒体文本的 ABSA 研究,同时提供原始和清洗版本以支持鲁棒性研究和实际应用开发。
建议
- 禁止将数据集用于违反 Twitter 开发者协议或适用隐私法的任何目的。
- 基于该数据集发表研究成果时,必须引用相关论文。




