遇见数据集

Evaluating AI agent permission caps under prompt injection with reference scoring and enforced execution — data and code

收藏
Zenodo2026-09-27 更新2026-10-01 收录
官方服务:

资源简介:

Data and code release 2.0.0 supporting the revised manuscript Evaluating AI agent permission caps under prompt injection with reference scoring and enforced execution. This update adds the corrected authority mapping and reference replay, enforced-cap and repeat-run materials, independent coder data, primary-policy uncertainty analyses, and a matched offline audit of attack success and benign completion. The revision-analysis folder contains reproducible inputs, scripts, full-precision results and hashes; original-study preserves earlier materials explicitly labelled as legacy. The paired audit reproduces S11, including the GPT-4o-mini A4 ASR interval of -1.91 to 5.88 percentage points. Across seven selected model-policy conditions, maximum absolute point discrepancies are 1.91 points for attack success and 7.22 for benign completion; these are not error bounds. No new model calls were made for this revision. The README documents offline reproduction, its scope, and upstream data and licence information. All attacks use mock AgentDojo environments. This is a research-material release, not a claim of journal acceptance or a validated deployment policy.

提供机构:
Zenodo
创建时间:
2026-09-27
二维码
社区交流群
二维码
科研交流群
商业服务