Algerian Dialect Dataset
收藏资源简介:
阿尔及利亚方言数据集是由阿卜杜勒哈米德·梅赫里大学团队构建的大规模情感标注语料库,包含45,000条来自30余个阿尔及利亚媒体频道YouTube评论的方言文本。数据集采用五级情感标注体系(从非常消极到非常积极),完整保留了包括表情符号、代码转换等真实网络语言特征,并附带发布时间、点赞数等元数据。通过严格的母语者人工标注流程,该资源有效填补了阿拉伯方言NLP研究空白,适用于情感分析模型训练、社会舆情研究及跨方言迁移学习等场景。
The Algerian Dialect Dataset is a large-scale sentiment-annotated corpus constructed by a research team from Abdelhamid Mehri University. It contains 45,000 dialectal texts sourced from YouTube comments across more than 30 Algerian media channels. The dataset adopts a five-level sentiment annotation system ranging from "very negative" to "very positive", fully preserves authentic internet language features including emojis and code-switching, and is accompanied by metadata such as publication time and like counts. Through a strict native speaker manual annotation workflow, this resource effectively fills the research gap in Arabic dialect natural language processing (NLP), and is suitable for scenarios including sentiment analysis model training, social public opinion research, and cross-dialect transfer learning.




