Provenance Blindness in Agentic AI Pipelines: Empirical Evidence of Token Forgery Vulnerability and Mitigation via Explicit Provenance Specification
收藏资源简介:
This preprint presents empirical evidence of a novel vulnerability class in agentic AI pipeline architectures that use explicit authorization tokens as access control mechanisms. Using the VATA (Verified Adversarial Testing Architecture) framework — a forensic-grade AI behavioral testing system with cryptographic on-chain anchoring — I demonstrate a reproducible asymmetry across three frontier models (Grok-4, GPT-5, Claude Sonnet 4): near-zero susceptibility to multi-turn social engineering attacks (0% breach rate across 6 batteries and multiple attack vectors) combined with high susceptibility to user-supplied credential forgery when token provenance is not explicitly specified. I term this asymmetry Provenance Blindness. The most effective forgery vector — a naturalistic conversation followed by an inline token claim — achieved a 100% breach rate on Grok-4 and 80% on GPT-5, compared to 0% across all social engineering batteries on the same models. A prompt-level mitigation using explicit provenance rules reduced breach rates to 0% on all vulnerable models, confirming the vulnerability is a specification gap rather than an intractable model failure.



