VATA BEC Results
收藏资源简介:
This paper documents twenty-two batteries (B90–B111) of the VATA adversarial testing framework investigating Business Email Compromise (BEC) fraud detection capabilities and failures across six frontier AI models. We identify a complete failure mechanism taxonomy — five distinct failure profiles across six models — and develop deployable prompt-level fixes that restore correct fraud detection to near-100% accuracy. We confirm that FIX1, the primary payroll diversion fix, is adversarially robust: it cannot be overridden by adversarial system prompt injection across any of the six tested models. A novel supply chain compromise scenario not previously documented in any VATA battery is flagged correctly by all six models without instruction, confirming that resistance generalizes beyond known patterns. A multi-scenario stress test (B111) across eight distinct BEC attack categories reveals a previously undocumented universal blind spot: HR benefits beneficiary change fraud is approved by Claude, GPT-4o, GPT-5.4, and Grok at 0% correct detection rates. All findings are SHA-256 hashed and anchored to Ethereum Mainnet before disclosure.



