COLDATASET
收藏资源简介:
COLDATASET是由清华大学智能技术与系统国家重点实验室开发的,包含37,480条中文评论数据集,旨在分析和检测中文中的攻击性语言。数据集覆盖了种族、性别和地区等多个敏感话题,每条评论都标有是否具有攻击性的二元标签。创建过程中,数据通过关键词查询和相关子话题爬虫从社交媒体平台收集,经过人工标注和模型辅助筛选,确保数据质量和相关性。COLDATASET的应用领域主要集中在提升社交媒体平台的文明程度和部署预训练语言模型的安全性,解决网络环境中的语言攻击问题。
COLDATASET is developed by the State Key Laboratory of Intelligent Technology and Systems, Tsinghua University. It encompasses 37,480 Chinese comment samples, aiming to analyze and detect offensive language in Chinese. The dataset covers multiple sensitive topics such as race, gender and region, and each comment is labeled with a binary tag indicating whether it is offensive. During its creation, data was collected from social media platforms via keyword queries and crawlers for relevant sub-topics, followed by manual annotation and model-assisted filtering to ensure data quality and relevance. The application scenarios of COLDATASET mainly focus on improving the civility of social media platforms and enhancing the safety of pre-trained language models, so as to address the problem of language attacks in online environments.

- 1COLD: A Benchmark for Chinese Offensive Language Detection清华大学智能技术与系统国家重点实验室 · 2022年



