Boiling the Frog
收藏资源简介:
Boiling the Frog是由伊卡洛斯基金会、罗马大学等多机构联合构建的基准测试数据集,旨在评估企业环境中工具型AI代理对渐进式攻击的脆弱性。该数据集包含157条多轮交互链,每条链包含1至20轮用户与AI代理的对话,模拟真实办公场景中的文件编辑操作,数据来源于人工设计的风险场景,并映射至欧盟《人工智能法案》的高风险情境。数据集通过沙盒化Docker工作空间实现,聚焦于操作风险,用于测试AI代理在持久化工作空间中是否会将文件状态修改为不安全配置,从而解决从响应风险向操作风险转变的安全评估难题。
Boiling the Frog is a benchmark dataset jointly constructed by multiple institutions including the Icarus Foundation and Sapienza University of Rome, aiming to evaluate the vulnerability of tool-augmented AI Agents against progressive attacks in enterprise environments. This dataset comprises 157 multi-turn interaction chains, with each chain containing 1 to 20 rounds of conversations between users and AI Agents, simulating file editing operations in real office scenarios. The data is sourced from manually designed risk scenarios and mapped to high-risk scenarios stipulated in the EU Artificial Intelligence Act. Implemented via sandboxed Docker workspaces, this dataset focuses on operational risks, and is designed to test whether AI Agents will alter file states to unsafe configurations in persistent workspaces, thereby resolving the security evaluation challenge brought about by the transition from response risks to operational risks.
数据集详情总结
标题: Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety
作者: Piercosma Bisconti, Matteo Prandi, Federico Pierucci, 等共14位作者
版本信息: arXiv:2605.22643v2,提交于2026年5月21日(v1),2026年5月22日更新(v2,当前版本)
所属领域: 计算机科学 > 计算与语言 (cs.CL)
核心贡献
该论文引入了一个名为 Boiling the Frog 的基准测试,专门用于评估在企业和办公环境中部署的、使用工具的AI模型面对渐进式攻击时的安全性。
评估方法
- 评估对象: 具备工具使用能力的AI代理模型。
- 场景设计: 每个测试场景从良性的工作空间编辑开始,随后逐步引入包含风险的请求。基准测试聚焦于有状态的多轮评估:场景链暴露一个持久的工作空间,将风险载荷放置在对话序列中的受控位置,并根据最终产生的工件状态是否变得不安全来进行评分。
- 风险分类: 场景基于一个三级操作风险分类法,该分类法根植于“温水煮青蛙”风险、AI法案附件I和附件III的高风险情境,以及欧盟AI法案的通用人工智能(GPAI)实践准则。
主要结果
- 总体攻击成功率(ASR): 在9个模型的测试面板中,严格攻击成功率为 44.4%。
- 模型级ASR范围: 从 20.5%(Claude Haiku 4.5)到 92.9%(Gemini 3.1 Flash Lite),Seed 2.0 Lite的ASR也超过80%。
- 场景类别级ASR: 在“实践准则”的失控场景中,平均攻击成功率高达 93.3%。

- 1Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety伊卡洛斯基金会; 罗马大学; 圣安娜高等研究学院; 同济大学法学院; AIQI联盟; BeEthical.be; 天主教圣心大学; 皮卡迪利实验室; 阿姆斯特丹自由大学; 独立研究者 · 2026年




