AfriHate
收藏资源简介:
AfriHate数据集是一个涵盖15种非洲语言的仇恨言论和侮辱性语言的多语言数据集,由多个非洲大学和研究机构合作创建。该数据集包含来自2012年至2023年的推文,每条推文均由熟悉当地文化的母语者进行标注,标注类别包括仇恨、侮辱/冒犯或中性。数据集的内容涉及种族、政治、性别、宗教等多个领域,旨在为研究社区提供高质量的数据基础,帮助开发针对非洲语言的仇恨言论检测工具。数据集的创建过程包括数据收集、预处理、语言识别和标注等步骤,特别关注了非洲语言的特殊性和文化背景。该数据集的应用领域包括自然语言处理、社交媒体内容审核以及非洲语言研究。
The AfriHate dataset is a multilingual dataset encompassing hate speech and offensive language in 15 African languages, collaboratively created by multiple African universities and research institutions. The dataset includes tweets from 2012 to 2023, with each tweet annotated by native speakers familiar with the local culture. The annotation categories include hate, offense/offensive, or neutral. The content of the dataset spans multiple domains such as race, politics, gender, and religion, aiming to provide the research community with a high-quality data foundation to develop hate speech detection tools for African languages. The creation process of the dataset includes data collection, preprocessing, language identification, and annotation, with a particular focus on the uniqueness and cultural context of African languages. The application areas of the dataset include natural language processing, social media content moderation, and research on African languages.




