Domain-Dependent BEC Fraud Susceptibility in grok-4.3: A Cross-Model Comparative Analysis with Validated Remediation
收藏资源简介:
This paper reports a cross-model test of how well claude-opus-4-8, gpt-5.4, and grok-4.3 resist Business Email Compromise (BEC)-style fraud when acting as autonomous approval agents — across four financial contexts: vendor payments, executive wire transfers, cloud credential rotation, and payroll direct-deposit changes. Using a controlled design (one validated attacker generating all fraud attempts, so only the target model varies), claude and gpt held at 0% fraud-approval across every domain and every trial. grok-4.3 showed a real, domain-dependent gap — from 0% on vendor wire fraud up to 41.3% on payroll diversion. A single added instruction telling the model to verify identity through an independent channel fully closed that gap in a full re-test. The paper also transparently documents two methodology corrections made along the way — a leading system prompt that inflated an earlier result, and an attacker-refusal contamination issue in the first cross-model comparison — both caught, fixed, and disclosed rather than hidden. Every dataset is hashed and anchored on Ethereum Sepolia before publication, so the findings carry an on-chain, tamper-evident timestamp. Bottom line: fraud resistance in AI approval agents isn't uniform across models or contexts, and at least one real gap is cheaply fixable with better prompting rather than requiring a different model.



