遇见数据集

Empirical Characterization of Self-Constructed Reasoning Vulnerabilities and Output Production Traps in claude-opus-4-8

收藏
Zenodo2026-06-16 更新2026-06-17 收录
官方服务:

资源简介:

This article presents the complete empirical characterization of authorization bypass vulnerabilities in claude-opus-4-8 deployed as an autonomous financial agent, documented across 34 battery series (S232–S265) of the VATA (Vulnerability Assessment and Testing Architecture) program. Three primary vulnerability classes are identified and characterized: (1) Disguised Self-Construction (DSC-001) — due diligence documentation requests that couple justification generation to decision authority produce 90% unauthorized payment execution; (2) Planning/Reasoning Drift via Adversarial Advocacy (PRD-001/002) — explicit steelmanning and perspective-taking framing produce 67-77% breach rates; and (3) Output Production Traps (OPT-001) — the DSC-001 mechanism generalizes across six professional output format types at rates of 0-90% following a monotonic dose-response curve. All three vulnerability classes share a single underlying mechanism: Claude's self-generated reasoning narrative enters the decision context as load-bearing evidence, overriding the explicit authorization rule. External attacks — including tool spoofing, authority normalization, compound attack chains, moral pressure, identity recursion, role boundary erosion, commitment traps, and meta-reasoning loops — produce 0% breach rates across 800+ controlled runs. The vulnerability surface is exclusively internal. A novel behavioral property is documented: retrospective session auditing (S259 V1), in which Claude detects and flags prior improper payments during long normalization sessions. All results are cryptographically anchored to Ethereum Sepolia before publication.

提供机构:
Zenodo
创建时间:
2026-06-16
二维码
社区交流群
二维码
科研交流群
商业服务