AIACW Teasers v1.0 — Model-Poisoning Cross-Model Sweep + Judge Ablation + Monte Carlo Defense Feasibility (Pharos Lighthouse Empirical Bundle)
收藏资源简介:
Empirical bundle supporting the Pharos Lighthouse programme (parent record: 10.5281/zenodo.19645912, UN CTED submission 2026-04-18). This version (v1.0) extends v0.3 along the threat-side pillar: THREAT SIDE — Model-Poisoning Teaser v1.0 • Cross-model T1 sweep: claude-opus-4-7 / claude-sonnet-4-6 / claude-haiku-4-5-20251001, identical 240-cell factorial per model • Judge-model ablation: opus run re-judged by haiku-as-judge; Cohen's κ = 0.646 on refusal labels, Pearson r = 0.789 on harm severity • Pre-registered Case A / Case B / Mixed framework • Headline: Sonnet → Case B, Haiku → Case B, Opus → Mixed (boundary case on Persona Override only; classification flips across runs and judges) DEFENSE SIDE — Global Defense Sim v0.1 (carry-forward from v0.2/v0.3) • 5000-episode Monte Carlo, 30 jurisdictions, S0/S1/S2 paired counterfactual • Δ(S2 − S0) = −1.13 [−1.18, −1.09]; 0/5000 capability_override • Parameter sweeps establish enforcement strength as dominant lever (~10× attribution success and peer influence) Companion papers: P7 — Residual Manipulability of Frontier Aligned Language Models … P8 — Enforcement, Not Attribution: An Empirical Feasibility Envelope … Both targeting NDSS 2027. Code: MIT. Data / text: CC-BY-4.0. Master seed: 20260424. Bootstrap N=10,000.



