HateXplain
收藏资源简介:
HateXplain是首个针对可解释仇恨言论检测的基准数据集,由印度理工学院卡拉格普尔分校和汉堡大学共同创建。该数据集包含20,148条来自Twitter和Gab的帖子,每条帖子都从三个不同角度进行标注:基本的三类分类(仇恨、攻击性或正常)、目标社区(帖子中仇恨言论/攻击性言论的受害者社区)以及理由(标注决策所依据的帖子部分)。数据集的创建过程涉及使用Amazon Mechanical Turk进行标注,确保了数据的质量和多样性。HateXplain数据集的应用领域主要集中在提高仇恨言论检测模型的解释性和减少对目标社区的意外偏见,为未来的仇恨言论研究提供了一个基础资源。
HateXplain is the first benchmark dataset for explainable hate speech detection, co-created by the Indian Institute of Technology Kharagpur and the University of Hamburg. This dataset contains 20,148 posts from Twitter and Gab. Each post is annotated from three distinct perspectives: a basic three-category classification (hate speech, offensive language, or normal), the target community (the victimized community targeted by hate or offensive speech in the post), and rationales (the specific segments of the post that serve as the basis for the annotators' labeling decisions). The dataset's creation process involved using Amazon Mechanical Turk for annotation, which ensures the quality and diversity of the data. The application fields of the HateXplain dataset mainly focus on improving the interpretability of hate speech detection models and reducing unintended biases against target communities, providing a foundational resource for future hate speech research.
- 1HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection印度理工学院卡拉格普尔分校 · 2022年



