five

Tunizi: Tunisian Arabizi Sentiment Analysis Dataset

收藏
NIAID Data Ecosystem2026-03-12 收录
下载链接:
https://zenodo.org/record/4271535
下载链接
链接失效反馈
官方服务:
资源简介:
Tunizi is the first 100% Tunisian Arabizi sentiment analysis dataset. Tunisian Arabizi is the representation of the tunisian dialect written in Latin characters and numbers rather than Arabic letters.We gathered comments from social media platforms that express sentiment about popular topics. For this purpose, we extracted 100k comments using public streaming APIs.  Tunizi was preprocessed by removing links, emoji symbols, and punctuations. The collected comments were manually annotated using an overall polarity:   positive (1), negative (-1) and neutral (0) class.   We divided the dataset into separate training, validation and test sets, with a ratio of 7:1:2  with a balanced split where the number of comments from positive class and negative class are almost the same.
创建时间:
2020-12-09
5,000+
优质数据集
54 个
任务类型
进入经典数据集
二维码
社区交流群

面向社区/商业的数据集话题

二维码
科研交流群

面向高校/科研机构的开源数据集话题

数据驱动未来

携手共赢发展

商业合作