遇见数据集

AgentDojo-PROV: A W3C PROV-O Corpus of LLM Agent Executions

收藏
Zenodo2026-06-30 更新2026-08-02 收录
官方服务:

资源简介:

A public corpus of W3C PROV-O conformant provenance graphs of large language model (LLM) agent executions, generated by instrumenting the AgentDojo prompt-injection benchmark (agentdojo 0.1.35; suite set v1.2.2: workspace, banking, travel, slack) with a purely observing capture layer.The corpus is generated with six agent models spanning proprietary and open-weight families: DeepSeek-V3.2 (deepseek-chat) OpenAI GPT-5-nano (gpt-5-nano) Google Gemini 2.5 Flash (gemini-2.5-flash) Meta Llama 3.3 70B (llama-3.3-70b-instruct) Alibaba Qwen2.5 72B (qwen2.5-72b-instruct) Z.ai GLM-4.6 (glm-4.6) Each agent provides 2,944 traces (97 benign + 949×3 under the direct, important_instructions, and injecagent attacks), for 17,664 traces in total.Each agent run is recorded as a lossless transcript (raw tool outputs, arguments, any model reasoning, errors, per-call timing, the (utility, security) task outcome, and injection ground truth), from which a W3C PROV-O conformant provenance graph is derived as a pure function — so any change of labelling or representation is an offline recompute, never a model re-run. Every released graph passes both structural validation and a PROV-JSON ⇄ PROV-O RDF round-trip isomorphism check.

提供机构:
Zenodo
创建时间:
2026-06-30
二维码
社区交流群
二维码
科研交流群
商业服务