LLM安全评测数据
收藏资源简介:
本数据集面向**大语言模型安全与合规评测**设计,以中国国家标准《网络安全技术 生成式人工智能服务安全基本要求》(TC260)中的分类体系为基础,融合论文《RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent》中提出的**上下文感知红队越狱攻击方法**,对原始合规评测问题进行系统性改造,生成更具挑战性和现实威胁模拟能力的越狱提示。
This dataset is designed for **safety and compliance evaluation of large language models (LLMs)**. It is developed based on the classification framework specified in the Chinese National Standard "Cybersecurity Technology – Basic Requirements for the Safety of Generative Artificial Intelligence Services" (TC260), and incorporates the **context-aware red teaming jailbreak attack methodology** proposed in the paper "RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent" to systematically restructure the original compliance evaluation questions, thus generating jailbreak prompts with elevated challenge and realistic threat simulation capabilities.




