遇见数据集

Downstream Pipeline Contamination

收藏
Zenodo2026-06-07 更新2026-06-12 收录
官方服务:

资源简介:

This paper answers a question no one had tested empirically: when an AI agent is compromised by batch contamination, does that compromise travel downstream and infect the next agent in the pipeline? The short answer is — it depends on the architecture. The research runs three controlled experiments across five frontier AI models from Anthropic, OpenAI, and xAI. The first experiment tests a two-agent pipeline where a document extraction agent feeds output directly to a payment execution agent. When the handoff is framed as pre-authorized pipeline output, batch contamination propagates field-wide — breach rates between 33% and 50% across four of five models. A single compromised agent becomes a weapon against every agent downstream. The second and third experiments add a third agent — a dedicated approval role sitting between procurement and payment. With that intermediate role present, contamination stops. Zero breaches across all five models. The third experiment then asks whether an attacker can restore propagation by framing every handoff as pipeline-verified and pre-cleared. The answer is no. Zero breaches again. The structural finding is that role differentiation — not framing controls — is the architectural security variable. A named intermediate agent with a distinct role identity causes downstream agents to engage independent verification behavior regardless of what the upstream framing claims. The intermediate agent doesn't even need to catch the contamination itself. Its presence alone changes how the final execution agent reasons about its inputs. Every result is cryptographically anchored to Ethereum Sepolia before disclosure. The scripts are public. The data is reproducible. This is the first published empirical characterization of pipeline contamination propagation in multi-agent AI systems — and it comes with receipts.

提供机构:
Zenodo
创建时间:
2026-06-07
二维码
社区交流群
二维码
科研交流群
商业服务