AgentDojo-PROV: A W3C PROV-O Corpus of LLM Agent Executions
收藏资源简介:
A public corpus of W3C PROV-O conformant provenance graphs of large language model (LLM) agent executions, generated by instrumenting the AgentDojo prompt-injection benchmark (agentdojo 0.1.35; suite set v1.2.2: workspace, banking, travel, slack) with a purely observing capture layer.The corpus is generated with six agent models spanning proprietary and open-weight families: DeepSeek-V3.2 (deepseek-chat) OpenAI GPT-5-nano (gpt-5-nano) Google Gemini 2.5 Flash (gemini-2.5-flash) Meta Llama 3.3 70B (llama-3.3-70b-instruct) Alibaba Qwen2.5 72B (qwen2.5-72b-instruct) Z.ai GLM-4.6 (glm-4.6) Each agent provides 2,944 traces (97 benign + 949×3 under the direct, important_instructions, and injecagent attacks), for 17,664 traces in total.Each agent run is recorded as a lossless transcript (raw tool outputs, arguments, any model reasoning, errors, per-call timing, the (utility, security) task outcome, and injection ground truth), from which a W3C PROV-O conformant provenance graph is derived as a pure function — so any change of labelling or representation is an offline recompute, never a model re-run. Every released graph passes both structural validation and a PROV-JSON ⇄ PROV-O RDF round-trip isomorphism check.



