SuperCLUE-Safety
收藏资源简介:
SuperCLUE-Safety是由中文自然语言处理与理解团队开发的多轮开放式问题对抗安全基准,用于评估大型语言模型在中文环境下的安全性。该数据集包含4912个开放式问题,覆盖超过20个安全子维度,旨在通过对抗性人机交互和对话,系统地评估模型的安全性。数据集设计考虑了传统安全、负责任AI和指令攻击的鲁棒性,以促进更安全、更可信赖的大型语言模型的发展。
SuperCLUE-Safety is a multi-turn open-ended question adversarial safety benchmark developed by the Chinese natural language processing and understanding team, designed to evaluate the safety of large language models in Chinese scenarios. This dataset contains 4912 open-ended questions covering more than 20 safety sub-dimensions, aiming to systematically assess model safety through adversarial human-machine interaction and dialogue. The dataset is constructed with consideration of robustness against traditional security risks, responsible AI principles and instruction attacks, to promote the development of safer and more trustworthy large language models.

- 1SC-Safety: A Multi-round Open-ended Question Adversarial Safety Benchmark for Large Language Models in Chinese中文自然语言处理与理解团队 · 2023年



