GMHP7k: A corpus of german misogynistic hatespeech posts

NIAID Data Ecosystem2026-05-02 收录

下载链接：

https://zenodo.org/record/10513452

下载链接

链接失效反馈

官方服务：

资源简介：

We provide a german corpus consisting of 7,061 posts authored by users of social media platforms. A group of volunteers annotated each post according to hatespeech and misogynistic/misogynous hatespeech in a binary fashion. The interrater reliability over all annotators according to Fleiss’ Kappa is 0.6409 for hatespeech and 0.8258 for misogynistic hatespeech. Furthermore, baseline measurements with machine learning based text classification with BERT are presented. Initial experiments with the corpus achieve macro average F1-scores up to 0.79 for hatespeech and 0.75 for misogynistic hatespeech. The dataset of the corpus on German Misogynistic Hatespeech Posts (GMHP7k) is publicly available.

创建时间：

2024-06-03

5,000+

优质数据集

54 个

任务类型

进入经典数据集