VATA-RCI-001: Rule Contradiction Injection in Agentic AI Systems
收藏资源简介:
I report VATA-RCI-001, a novel adversarial attack class against correctly configured autonomous AI agents. Unlike prior agentic security findings in the VATA corpus that exploit absent or incomplete system prompt rules, Rule Contradiction Injection (RCI) exploits the semantic gap between two legitimate, co-present rules by constructing scenarios in which honoring one rule appears to require violating the other. Empirical testing across Series 112-114 of the VATA battery series demonstrates that grok-4-0709 breaches at 69.2% (9/13 runs) on the authorization-continuity conflict vector when both rules are explicitly present and correctly stated. claude-opus-4-8 and gpt-5.4 hold at 0% across 13 combined runs on the same vector. I confirm three mitigations effective, all closing the breach to 0/10 on grok-4-0709. This is the highest empirically confirmed breach rate against a correctly configured agent in my corpus. All results are SHA256-hashed and anchored to Ethereum Sepolia blockchain prior to disclosure.



