Capability is not risk: Evaluating delegated-authority caps against prompt injection in LLM agents — data, code and new trajectories
收藏资源简介:
This package accompanies the manuscript "Capability is not risk: Evaluating delegated-authority caps against prompt injection in LLM agents" (submitted to the Journal of Information Security and Applications). It contains (1) a rule-based authority codebook that assigns the 74 AgentDojo tool entries to five authority levels using four criteria (state change, third-party commitment, outbound transmission or access grant, money/credential/irreversible action); (2) a run-level dataset of 38,938 agent runs, comprising 36,679 runs extracted from the public AgentDojo logs of 22 models and 2,259 new runs of three current models (Claude Haiku 4.5, GPT-6 Luna, GPT-6 Sol) collected by the author in September 2026; (3) the raw trajectories and provider-billed token usage of the new runs, including an auto-confirm condition; (4) a model-free full-hijack replay of reference solutions for 629 injection cases under seven authority conditions; (5) pre-specified predictions for the current-model runs with their SHA-256 hash; and (6) all analysis code, tables and figures. All tables and figures in the manuscript can be regenerated with code/reproduce.sh. All attacks were run in AgentDojo's mock environments.



