C-SafeQA
收藏资源简介:
C-SafeQA是一个面向中文安全评估的响应级基准数据集,由安徽火花盾智能科技、科大讯飞、北京大学、上海创新研究院和上海人工智能实验室联合创建,旨在严格区分查询风险与响应违规,为评估大语言模型输出安全性与自动化安全裁判的可靠性提供统一标尺。该数据集包含538条基础查询和8,877条对抗性查询,由四个全模型LLM分别回答生成总计37,660条查询-响应记录,每条记录经多模型共识投票与分层专家盲审获得安全、不安全或有争议的三类标签。数据集构建过程基于内部安全政策分解的269个风险点,人工撰写基础查询并应用21种中文对抗变换方法扩展,确保政策可追溯性与攻击多样性。C-SafeQA可用于评估目标LLM在多种风险类别和攻击条件下的安全表现,以及审计自动化安全裁判的召回率、误报率等指标,揭示安全评估中模型与裁判的双重局限性。
C-SafeQA is a response-level benchmark dataset for Chinese safety evaluation, jointly developed by Anhui SparkShield Intelligent Technology, iFLYTEK, Peking University, Shanghai Innovation Research Institute, and Shanghai AI Laboratory. It aims to strictly differentiate between query risks and response violations, providing a unified benchmark for evaluating the safety of large language model (LLM) outputs and the reliability of automated safety referees. This dataset includes 538 base queries and 8,877 adversarial queries. A total of 37,660 query-response pairs were generated by having four full-model LLMs independently answer each query. Each record is assigned one of three labels—safe, unsafe, or controversial—through multi-model consensus voting and hierarchical expert blind review. The dataset construction process is based on 269 risk points decomposed from internal safety policies. Base queries were manually authored and expanded using 21 Chinese adversarial transformation methods, ensuring policy traceability and attack diversity. C-SafeQA can be used to assess the safety performance of target LLMs across diverse risk categories and attack scenarios, as well as audit metrics of automated safety referees such as recall rate and false positive rate, uncovering the dual limitations of both models and referees in safety evaluation.

- 1Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges安徽火花盾智能科技; 科大讯飞; 北京大学; 上海创新研究院; 上海人工智能实验室 · 2026年



