遇见数据集

sathvikchandra77/afrisenti

收藏
Hugging Face2026-05-21 更新2026-05-31 收录
官方服务:

资源简介:

AfriSenti是最大的针对代表性不足的非洲语言的情感分析数据集,覆盖14种非洲语言(包括阿姆哈拉语、阿尔及利亚阿拉伯语、豪萨语、伊博语、基尼亚卢旺达语、摩洛哥阿拉伯语、莫桑比克葡萄牙语、尼日利亚皮钦语、奥罗莫语、斯瓦希里语、提格里尼亚语、特威语、齐聪加语和约鲁巴语),包含超过110,000条带注释的推文。该数据集用于首届以非洲为中心的SemEval共享任务(SemEval 2023任务12),支持情感分类、情感强度分析和情绪检测等多种任务,数据字段包括推文文本和情感标签(正面、负面、中性),并分为训练集、验证集和测试集。

AfriSenti is the largest sentiment analysis dataset for under-represented African languages, covering 110,000+ annotated tweets in 14 African languages (Amharic, Algerian Arabic, Hausa, Igbo, Kinyarwanda, Moroccan Arabic, Mozambican Portuguese, Nigerian Pidgin, Oromo, Swahili, Tigrinya, Twi, Xitsonga, and Yoruba). The dataset is used in the first Afrocentric SemEval shared task (SemEval 2023 Task 12) and supports various tasks such as sentiment classification, sentiment intensity analysis, and emotion detection, with data fields including tweet text and sentiment labels (positive, negative, neutral), and splits into train, validation, and test sets.

提供机构:
sathvikchandra77
二维码
社区交流群
二维码
科研交流群
商业服务