Ethnohate
收藏资源简介:
Ethnic hate speech is a form of intersectional violence that affects indigenous groups in Mexico. Despite the seriousness of this social phenomenon, there is a lack of computational resources, specifically labeled data sets, that allow the development of automated tools for its detection. Particularly in Mexico, ethnic hate speech has shown greater prevalence and impact, particularly directed at the country's 68 indigenous communities. To address this gap, we present EthnoHate, a tagging dataset for automatic detection of ethnic hate speech in Mexico. The annotation scheme includes three categories: Hate (speech that attacks, insults or dehumanizes indigenous people or communities due to their ethnic condition), No Hate (neutral, positive or merely informative content without hate) and Unrelated Hate (hate speech directed at another population, but related to ethnic origin).The corpus consists of 17,000 public posts extracted from X (formerly Twitter). The majority class is No Hate (53%), followed by Hate (43%), while Unrelated Hate (4%) is the least represented class. the dataset was divided into three subsets: 70% of the instances were allocated for model training, 10% were reserved for hyperparameter tuning (development), and 20% for final evaluation (testing). This partition was done randomly, maintaining the class distribution.



