NoisyHate
收藏资源简介:
NoisyHate: Mining and Curating Human-Written Perturbations for Benchmarking Content Moderation Models Online texts with toxic content are a clear threat to the users on social media in particular and society in general. Although many platforms have adopted various measures (e.g., machine learning-based hate-speech detection systems) to diminish their effect, toxic content writers have also attempted to evade such measures by using cleverly modified toxic words, so-called human-written text perturbations. Therefore, to help build automatic detection tools to recognize those perturbations, prior methods have developed sophisticated techniques to generate diverse adversarial samples. However, we note that these ``algorithms"-generated perturbations do not necessarily capture all the traits of ``human"-written perturbations. Therefore, in this paper, we introduce a novel, high-quality dataset of human-written perturbations, named as NoisyHate, that was created from real-life perturbations that are both written and verified by human-in-the-loop. We show that perturbations in NoisyHate have different characteristics than prior algorithm-generated toxic datasets show, and thus can be in particular useful to help develop better toxic speech detection solutions. We thoroughly validate NoisyHate against state-of-the-art language models, such as BERT and RoBERTa, and black box APIs, such as Perspective API, on two tasks, such as perturbation normalization and understanding.
NoisyHate:面向内容审核模型基准测试的人工编写扰动数据集挖掘与整理 带有恶意内容的在线文本,尤其对社交媒体用户乃至整个社会而言,均构成显著威胁。尽管诸多平台已采用多种措施(例如基于机器学习的仇恨言论检测系统)以削弱其负面影响,但恶意内容撰写者亦会通过精心修改恶意词汇——即所谓的人工编写文本扰动——来规避此类检测手段。 为此,为助力开发可识别此类扰动的自动检测工具,此前的研究已开发出多种复杂技术以生成多样化的对抗样本。然而,我们注意到,这些算法生成的扰动未必能完全涵盖人工编写扰动的全部特征。 为此,本文提出一款全新的高质量人工编写扰动数据集NoisyHate,其样本源自经人机回圈(human-in-the-loop)流程撰写并验证的真实扰动文本。我们的实验表明,NoisyHate中的扰动文本与此前算法生成的恶意数据集所呈现的特征存在显著差异,因此该数据集可有效助力开发更优异的恶意言论检测方案。我们通过两项任务——扰动文本归一化与扰动文本理解——,针对多款最先进语言模型(如BERT、RoBERTa)以及黑盒API(如Perspective API)对NoisyHate进行了全面验证。



