CHiSafetyBench
收藏资源简介:
该数据集是一个用于评估大型语言模型在中文学术环境中识别风险内容并拒绝回答风险问题的安全性能基准。它涵盖了一个分层的中文安全分类体系,该体系包含5个风险领域和31个类别。该数据集包含了多项选择题和问答任务,旨在衡量大型语言模型在中文环境下的安全处理能力。数据集规模为2323个基准样本,任务重点在于风险内容的识别以及对于风险问题的拒绝回答。
This dataset is a safety performance benchmark for evaluating large language models' capacity to identify risky content and decline to answer risky questions within the Chinese academic context. It encompasses a hierarchical Chinese safety classification system, which comprises 5 risk domains and 31 categories. The dataset contains multiple-choice questions and question-answering tasks, aiming to measure the safety handling capabilities of large language models in the Chinese context. The dataset consists of 2,323 benchmark samples, with the tasks focusing on the identification of risky content and the refusal to respond to risky questions.




