智能体安全评测数据集
收藏资源简介:
本数据集聚焦AI智能体的安全防御能力评测,覆盖提示注入、越狱攻击、对抗性后缀、编码混淆、隐私提取、拒绝服务等主流攻击类型。数据内容包括评测任务元数据、攻击样本(含载荷及目标)、智能体响应(输出内容、是否拒绝、延迟)、防御检测结果(是否触发、防护动作、是否绕过)及安全评测指标(攻击成功率、防御拒绝率、安全评分、鲁棒性等级)。适用于智能体安全能力评估、防御策略对比选型、红蓝对抗演练、安全微调效果验证、新型攻击防御研究及合规审计等场景。
This dataset focuses on the safety and defense capability evaluation of AI Agents, covering mainstream attack types including Prompt Injection, Jailbreak Attack, Adversarial Suffix, Encoding Obfuscation, Privacy Extraction, Denial of Service (DoS), etc. The data content includes metadata of evaluation tasks, attack samples (including payloads and targets), agent responses (output content, rejection status, latency), defense detection results (whether the attack is triggered, protective actions, whether the attack is bypassed), and safety evaluation metrics (attack success rate, defense rejection rate, safety score, robustness level). It is applicable to scenarios such as agent security capability evaluation, defense strategy comparison and selection, red-blue confrontation drills, safety fine-tuning effect verification, new attack and defense research, and compliance audits.




