遇见数据集

GMHP7k: A corpus of german misogynistic hatespeech posts

收藏
Zenodo2024-03-29 更新2026-05-26 收录
官方服务:

资源简介:

We provide a german corpus consisting of 7,061 posts authored by users of social media platforms. A group of volunteers annotated each post according to hatespeech and misogynistic/misogynous hatespeech in a binary fashion. The interrater reliability over all annotators according to Fleiss’ Kappa is 0.6409 for hatespeech and 0.8258 for misogynistic hatespeech. Furthermore, baseline measurements with machine learning based text classification with BERT are presented. Initial experiments with the corpus achieve macro average F1-scores up to 0.79 for hatespeech and 0.75 for misogynistic hatespeech. The dataset of the corpus on German Misogynistic Hatespeech Posts (GMHP7k) is publicly available.

本研究提供一则由社交媒体平台用户发布的7061条帖文组成的德语语料库。一组志愿者依据仇恨言论(hatespeech)与厌女仇恨言论(misogynistic hatespeech)的二元分类标准,对每条帖文完成标注。经弗莱伊斯kappa系数(Fleiss’ Kappa)测算,全体标注者在仇恨言论标注任务上的评分者间信度为0.6409,在厌女仇恨言论标注任务上为0.8258。此外,本研究还给出了基于机器学习的文本分类模型BERT的基准测试结果。使用该语料库开展的初步实验显示,仇恨言论分类任务的宏平均F1值最高可达0.79,厌女仇恨言论分类任务则为0.75。本德语厌女仇恨言论帖文语料库(GMHP7k)的数据集已公开可获取。

提供机构:
Zenodo
创建时间:
2024-03-29
二维码
社区交流群
二维码
科研交流群
商业服务