遇见数据集

Ethnohate

收藏
Zenodo2026-05-20 更新2026-05-26 收录
官方服务:

资源简介:

Ethnic hate speech is a form of intersectional violence that affects indigenous groups in Mexico. Despite the seriousness of this social phenomenon, there is a lack of computational resources, specifically labeled data sets, that allow the development of automated tools for its detection. Particularly in Mexico, ethnic hate speech has shown greater prevalence and impact, particularly directed at the country's 68 indigenous communities. To address this gap, we present EthnoHate, a tagging dataset for automatic detection of ethnic hate speech in Mexico. The annotation scheme includes three categories: Hate (speech that attacks, insults or dehumanizes indigenous people or communities due to their ethnic condition), No Hate (neutral, positive or merely informative content without hate) and Unrelated Hate (hate speech directed at another population, but related to ethnic origin).The corpus consists of 17,000 public posts extracted from X (formerly Twitter). The majority class is No Hate (53%), followed by Hate (43%), while Unrelated Hate (4%) is the least represented class. the dataset was divided into three subsets: 70% of the instances were allocated for model training, 10% were reserved for hyperparameter tuning (development), and 20% for final evaluation (testing). This partition was done randomly, maintaining the class distribution.

族群仇恨言论是一种交叉性暴力形式,对墨西哥的原住民群体造成侵害。尽管这一社会现象危害深重,但目前仍缺乏相应的计算资源——尤其是标注数据集——以开发用于自动检测此类言论的自动化工具。在墨西哥,族群仇恨言论的传播范围与影响尤为突出,其针对的对象正是该国68个原住民社区。为填补这一研究空白,我们提出了EthnoHate数据集,一款用于自动检测墨西哥境内族群仇恨言论的标注数据集。该数据集的标注体系包含三大类别:仇恨类(Hate),即因族群身份攻击、侮辱或非人化原住民个人或社区的言论;非仇恨类(No Hate),即不含仇恨倾向的中立、积极或单纯资讯性内容;无关仇恨类(Unrelated Hate),即针对其他群体的仇恨言论,但与族群出身相关。该语料库包含从X(原Twitter)平台抓取的17000条公开帖子。数据集中占比最高的类别为非仇恨类(53%),其次为仇恨类(43%),占比最低的无关仇恨类仅占4%。本数据集被随机划分为三个子集:70%的样本用于模型训练,10%预留用于超参数调优(开发集),剩余20%用于最终模型评估(测试集),且划分过程严格保留了原有的类别分布。

提供机构:
Zenodo
创建时间:
2026-05-20
二维码
社区交流群
二维码
科研交流群
商业服务