遇见数据集

Dynamically-Generated-Hate-Speech-Dataset

收藏
OpenML2025-02-23 更新2025-12-20 收录
官方服务:

资源简介:

The Dynamically Generated Hate Speech Dataset is provided in one table. 'acl.id' is the unique ID of the entry. 'Text' is the content which has been entered. All content is synthetic. 'Label' is a binary variable, indicating whether or not the content has been identified as hateful. It takes two values: hate, nothate. 'Type' is a categorical variable, providing a secondary label for hateful content. For hate it can take five values: Animosity, Derogation, Dehumanization, Threatening and Support for Hateful Entities. Please see the paper for more detail. For nothate the 'type' is 'none'. In round 1 the 'type' was not given and is marked as 'notgiven'. 'Target' is a categorical variable, providing the group that is attacked by the hate. It can include intersectional characteristics and multiple groups can be identified. For nothate the type is 'none'. Note that in round 1 the 'target' was not given and is marked as 'notgiven'. 'Level' reports whether the entry is original content or a perturbation. 'Round' is a categorical variable. It gives the round of data entry (1, 2, 3 or 4) with a letter for whether the entry is original content ('a') or a perturbation ('b'). Perturbations were not made for round 1. 'Round.base' is a categorical variable. It gives the round of data entry, indicated with just a number (1, 2, 3 or 4). 'Split' is a categorical variable. it gives the data split that the entry has been assigned to. This can take the values 'train', 'dev' and 'test'. The choice of splits is explained in the paper. 'Annotator' is a categorical variable. It gives the annotator who entered the content. Annotator IDs are random alphanumeric strings. There are 20 annotators in the dataset. 'acl.id.matched' is the ID of the matched entry, connecting the original (given in 'acl.id') and the perturbed version. paper_url = "https://aclanthology.org/2021.acl-long.132.pdf" original_data_url = "https://github.com/bvidgen/Dynamically-Generated-Hate-Speech-Dataset"

创建时间:
2025-02-23
二维码
社区交流群
二维码
科研交流群
商业服务