Paparare/toxic_benchmark_2024
收藏资源简介:
Paparare/toxic_benchmark_2024是由哥伦比亚大学Xinmeng Hou创建的一个用于检测有毒语言的综合性标注基准数据集。该数据集包含1942条数据,旨在通过人文研究为基础的规范性标注框架,确保对冒犯性语言的一致且无偏见的标注。数据集的创建过程结合了多源语言模型标注数据,通过小规模统计分析和实验验证了其有效性。该数据集主要应用于自然语言处理领域,旨在解决语言多样性保护和偏见减少的问题,特别是在非主流和非标准语言使用中。
Paparare/toxic_benchmark_2024 is a comprehensive annotated benchmark dataset for toxic language detection, created by Xinmeng Hou from Columbia University. Comprising 1942 instances, this dataset is designed to ensure consistent and unbiased annotation of offensive language via a humanistic research-based normative annotation framework. The dataset’s creation process incorporates multi-source language model annotated data, and its effectiveness has been validated via small-scale statistical analysis and experiments. This dataset is primarily applied in the field of natural language processing, aiming to address issues of language diversity preservation and bias mitigation, especially in the use of non-mainstream and non-standard languages.

- 1Mitigating Biases to Embrace Diversity: A Comprehensive Annotation Benchmark for Toxic Language哥伦比亚大学 · 2024年



