遇见数据集

ZHateBench: A Comprehensive Chinese Offensive Language Dataset with Harmful–Safe Pairs

收藏
Zenodo2025-08-18 更新2026-05-26 收录
官方服务:

资源简介:

ZHateBench is a large-scale, generation-based dataset for Chinese offensive language detection, consisting of over 53,609 samples across three major categories: sexual content, abusive language, and social bias. Each entry is presented as a Harmful–Safe pair, generated via LLMs using carefully designed prompts and keyword control. The dataset supports three subtypes: SexHarmSet: sexual and suggestive language AbuseSet: insults, profanity, and personal attacks BiasSet: including gender, occupation, region, and ethnicity-related biases Our huggingface: https://huggingface.co/datasets/RYOAL/ZHateBench Our github: https://github.com/royal12646/Chinese-offensive-language-detect ⚠️ Dataset Usage Disclaimer This dataset is created for research purposes in Chinese offensive language detection. It contains a wide range of potentially harmful, offensive, abusive, pornographic, or discriminatory text samples, which are included solely for the development, evaluation, and safety alignment of natural language processing (NLP) systems. Important Notes: All harmful content included in this dataset is either synthetically generated or collected from publicly available sources. It does not reflect the views or opinions of the author(s) and should not be interpreted as promoting or endorsing any form of hate, discrimination, or offensive behavior. This dataset is strictly prohibited from being used for any non-research purposes, including but not limited to propaganda, harassment, incitement, or real-world application in hostile contexts. By accessing or using this dataset, you acknowledge and accept the risks associated with the content, and agree to use it only for academic or engineering research aimed at improving model robustness, fairness, and safety. If any concerns arise regarding the ethical implications of this dataset, please contact the maintainer(s) for further discussion or removal requests. This dataset is intended to promote responsible AI development and support the creation of safer and more trustworthy language models.

提供机构:
Zenodo
创建时间:
2025-08-12
二维码
社区交流群
二维码
科研交流群
商业服务