ZHateBench: A Comprehensive Chinese Offensive Language Dataset with Harmful–Safe Pairs
收藏资源简介:
ZHateBench is a large-scale, generation-based dataset for Chinese offensive language detection, consisting of over 53,609 samples across three major categories: sexual content, abusive language, and social bias. Each entry is presented as a Harmful–Safe pair, generated via LLMs using carefully designed prompts and keyword control. The dataset supports three subtypes: SexHarmSet: sexual and suggestive language AbuseSet: insults, profanity, and personal attacks BiasSet: including gender, occupation, region, and ethnicity-related biases Our huggingface: https://huggingface.co/datasets/RYOAL/ZHateBench Our github: https://github.com/royal12646/Chinese-offensive-language-detect ⚠️ Dataset Usage Disclaimer This dataset is created for research purposes in Chinese offensive language detection. It contains a wide range of potentially harmful, offensive, abusive, pornographic, or discriminatory text samples, which are included solely for the development, evaluation, and safety alignment of natural language processing (NLP) systems. Important Notes: All harmful content included in this dataset is either synthetically generated or collected from publicly available sources. It does not reflect the views or opinions of the author(s) and should not be interpreted as promoting or endorsing any form of hate, discrimination, or offensive behavior. This dataset is strictly prohibited from being used for any non-research purposes, including but not limited to propaganda, harassment, incitement, or real-world application in hostile contexts. By accessing or using this dataset, you acknowledge and accept the risks associated with the content, and agree to use it only for academic or engineering research aimed at improving model robustness, fairness, and safety. If any concerns arise regarding the ethical implications of this dataset, please contact the maintainer(s) for further discussion or removal requests. This dataset is intended to promote responsible AI development and support the creation of safer and more trustworthy language models.
ZHateBench是一款大规模、基于生成式的中文攻击性语言检测数据集,共收录超53609条样本,涵盖三大核心类别:色情内容、辱骂性语言与社会偏见。每条数据均以“有害-安全”(Harmful–Safe)配对形式呈现,由大语言模型(LLMs)通过精心设计的提示词与关键词控制生成。 本数据集支持三类子数据集: SexHarmSet:包含色情与暗示性语言 AbuseSet:涵盖侮辱性言论、亵渎性语言及人身攻击 BiasSet:包含性别、职业、地域与族裔相关的偏见内容 我们的Hugging Face仓库:https://huggingface.co/datasets/RYOAL/ZHateBench 我们的GitHub仓库:https://github.com/royal12646/Chinese-offensive-language-detect ⚠️ 数据集使用免责声明 本数据集仅用于中文攻击性语言检测领域的研究工作。其包含大量潜在有害、攻击性、辱骂性、色情或歧视性文本样本,仅用于自然语言处理(NLP)系统的开发、评估与安全对齐任务。 重要提示: 1. 本数据集收录的所有有害内容均为人工合成生成或从公开渠道收集而来,并不代表作者的观点与立场,不应被解读为支持或宣扬任何形式的仇恨、歧视或攻击性行为。 2. 本数据集严格禁止用于任何非研究用途,包括但不限于宣传、骚扰、煽动或在敌对场景下的实际应用。 3. 访问或使用本数据集即代表您认可并承担相关内容带来的风险,且同意仅将其用于旨在提升模型鲁棒性、公平性与安全性的学术或工程研究。 4. 若您对本数据集的伦理影响存在任何疑虑,请联系数据集维护者进行进一步沟通或申请移除相关内容。 本数据集旨在推动负责任的人工智能开发,助力构建更安全、更可信的语言模型。



