Setswana Sentiment Dataset
收藏资源简介:
该Setswana情感数据集是由比勒陀利亚大学等机构构建的一个低资源语言标注资源,专门针对南非官方语言Setswana的推特文本进行情感分析。数据集包含3,565条推文,由三位母语标注者进行多批次人工标注,涵盖积极、消极、中性等五类情感标签,其中有效分类数据达3,454条,数据来源于2021年至2022年间的公开推特API,并经过语言识别和匿名化处理。创建过程采用LightTag标注工具,分七批次异步独立完成,并记录了每次标注的时间戳元数据,以支持质量审计。该数据集旨在解决非洲语言NLP资源稀缺问题,通过分析标注质量随时间下降的规律,为情感分类模型提供训练基准,并推动标注活动设计优化以提升低资源语言数据集的可靠性。
This Setswana sentiment dataset is a low-resource language annotation resource developed by the University of Pretoria and other institutions, specifically designed for sentiment analysis of Twitter texts in Setswana, an official language of South Africa. Comprising 3,565 tweets, the dataset was manually annotated in multiple batches by three native annotators, covering five sentiment labels including positive, negative, neutral and other categories, with 3,454 valid classified samples. The data was collected via the public Twitter API between 2021 and 2022, and underwent language identification and anonymization processing. Annotation was completed asynchronously and independently in seven batches using the LightTag annotation tool, and timestamp metadata for each annotation session was recorded to support quality auditing. This dataset aims to address the scarcity of NLP resources for African languages, serve as a training benchmark for sentiment classification models by analyzing the pattern of declining annotation quality over time, and promote the optimization of annotation activity design to enhance the reliability of low-resource language datasets.

- 1Temporal Simultaneity Predicts Annotation Quality in Sentiment Corpora比勒陀利亚大学·社会影响数据科学; 比勒陀利亚大学·非洲语言系; 帝国理工学院; 国立理工学院 · 2026年



