JailBench
收藏资源简介:
JailBench是由北京邮电大学可信分布式计算与服务教育部重点实验室创建的一个全面的中文安全性评估基准,旨在评估大型语言模型(LLM)在中文语境下的深层安全漏洞。该数据集包含10,800个查询,涵盖了5个不同的风险领域和40种具体的风险类型,通过自动化的数据扩展方法和创新的自动越狱提示工程框架(AJPE),提高了评估的有效性和效率。JailBench可以广泛应用于大型语言模型的安全性评估,特别是在中文语境下,有助于揭示模型的安全性和可信度方面的改进空间。
JailBench is a comprehensive Chinese safety evaluation benchmark developed by the Key Laboratory of Trusted Distributed Computing and Services of the Ministry of Education, Beijing University of Posts and Telecommunications. It is designed to assess the deep-seated security vulnerabilities of large language models (LLMs) within the Chinese context. This dataset contains 10,800 queries covering 5 distinct risk domains and 40 specific risk types. Through automated data expansion methods and an innovative automatic jailbreak prompt engineering framework (AJPE), it improves the effectiveness and efficiency of the evaluation process. JailBench can be widely applied to the safety evaluation of large language models, particularly in the Chinese context, and helps reveal the potential areas for improving model security and credibility.

- 1JailBench: A Comprehensive Chinese Security Assessment Benchmark for Large Language Models北京邮电大学可信分布式计算与服务教育部重点实验室 · 2025年



