Hate Speech and Offensive Language
收藏资源简介:
HSOL 是用于仇恨言论检测的数据集。作者从仇恨言论词典开始,其中包含被互联网用户识别为仇恨言论的单词和短语,由 Hatebase.org 编译。他们使用 Twitter API 搜索包含词典中术语的推文,从而产生了来自 33,458 位 Twitter 用户的推文样本。他们为每个用户提取了时间线,产生了一组 8540 万条推文。他们从这个语料库中随机抽取了 25k 条推文样本,其中包含词典中的术语,并由 CrowdFlower (CF) 工作人员手动编码。工人们被要求将每条推文标记为以下三类之一:仇恨言论、冒犯性但非仇恨言论或既非冒犯性又非仇恨言论。
HSOL is a dataset for hate speech detection. The dataset construction begins with a hate speech lexicon compiled by Hatebase.org, which contains words and phrases recognized as hate speech by internet users. They utilized the Twitter API to search for tweets containing terms from this lexicon, resulting in a sample of tweets from 33,458 unique Twitter users. They extracted the full tweet timelines for each of these users, yielding a corpus of 85.4 million tweets. They then randomly sampled 25,000 tweets containing lexicon terms from this corpus, which were manually annotated by CrowdFlower (CF) workers. Workers were instructed to label each tweet into one of three distinct categories: hate speech, offensive but not hate speech, and neither offensive nor hate speech.

- 首次发表了Hate Speech and Offensive Language数据集,该数据集由Zeerak Waseem和Dhruv Kulkarni创建,旨在识别和分类社交媒体中的仇恨言论和冒犯性语言。
- 数据集在多个自然语言处理和机器学习研究中被广泛应用,成为评估仇恨言论检测算法的标准基准之一。
- 随着数据集的普及,研究者们开始探索更复杂的模型和方法,以提高仇恨言论检测的准确性和鲁棒性。
- 数据集被用于多个国际会议和研讨会,促进了跨学科的合作和研究,特别是在社交媒体内容监管和伦理方面。
- 研究者们开始关注数据集的局限性和偏见问题,提出了改进和扩展数据集的建议,以更好地反映全球多样性和文化差异。



