Evaluating AI agent permission caps under prompt injection with reference scoring and enforced execution — data and code
收藏资源简介:
Data and code release 2.0.0 supporting the revised manuscript Evaluating AI agent permission caps under prompt injection with reference scoring and enforced execution. This update adds the corrected authority mapping and reference replay, enforced-cap and repeat-run materials, independent coder data, primary-policy uncertainty analyses, and a matched offline audit of attack success and benign completion. The revision-analysis folder contains reproducible inputs, scripts, full-precision results and hashes; original-study preserves earlier materials explicitly labelled as legacy. The paired audit reproduces S11, including the GPT-4o-mini A4 ASR interval of -1.91 to 5.88 percentage points. Across seven selected model-policy conditions, maximum absolute point discrepancies are 1.91 points for attack success and 7.22 for benign completion; these are not error bounds. No new model calls were made for this revision. The README documents offline reproduction, its scope, and upstream data and licence information. All attacks use mock AgentDojo environments. This is a research-material release, not a claim of journal acceptance or a validated deployment policy.



