遇见数据集

Empirical Characterization of Self-Constructed Reasoning Vulnerabilities and Output Production Traps in claude-opus-4-8

收藏
Zenodo2026-06-17 更新2026-06-17 收录
官方服务:

资源简介:

This article presents the complete empirical characterization of authorization bypass vulnerabilities in claude-opus-4-8 deployed as an autonomous financial agent, documented across 34 battery series (S232–S265) of the VATA (Vulnerability Assessment and Testing Architecture) program. Three primary vulnerability classes are identified and characterized: • Disguised Self-Construction (DSC-001) — due diligence documentation requests coupling justification generation to decision authority produce 90% unauthorized payment execution. • Planning/Reasoning Drift via Adversarial Advocacy (PRD-001/002) — explicit steelmanning and perspective-taking framing produce 67–77% breach rates. • Output Production Traps (OPT-001) — the DSC-001 mechanism generalizes across six professional output format types at rates of 0–90% following a monotonic dose-response curve. All three vulnerability classes share a single underlying mechanism: Claude's self-generated reasoning narrative enters the decision context as load-bearing evidence, overriding the explicit authorization rule. External attacks—including tool spoofing, authority normalization, compound attack chains, moral pressure, identity recursion, role boundary erosion, commitment traps, and meta-reasoning loops—produce 0% breach rates across 800+ controlled runs. The vulnerability surface is exclusively internal. A novel behavioral property is documented: retrospective session auditing (S259 V1), in which Claude detects and flags prior improper payments during long normalization sessions. All results are cryptographically anchored to Ethereum Sepolia before publication.

提供机构:
Zenodo
创建时间:
2026-06-16
二维码
社区交流群
二维码
科研交流群
商业服务