遇见数据集

AIACW Empirical Teasers v0.1 — Model-Level Attack Surface and Global-Defense Feasibility Supplements to Pharos Lighthouse

收藏
Zenodo2026-04-25 更新2026-05-26 收录
官方服务:

资源简介:

This deposit contains two minimum-viable empirical teasers that supplement the Pharos Lighthouse proposal (Zenodo DOI 10.5281/zenodo.19645912, UN CTED submission 2026-04-18). Pharos Lighthouse argues that AI-Assisted Cyber Weaponization is a society-scale risk requiring coordinated global defense. Reviewers have raised two empirical objections: (1) is model-level poisoning a real attack surface, or a rhetorical one? (2) is the proposed global-defense architecture feasible and effective, or speculative? This deposit contains the v0.1 teasers that answer each objection at the smallest responsible scale, with full code, initial parameters, random seeds, and pre-registered interpretation frameworks included so that no claim in the results is unfalsifiable. **Teaser 1 — Model-Poisoning Teaser v0.1 (threat side).**A pre-registered red-team evaluation of claude-opus-4-7 under five published adversarial-prompt attack classes (persona override; few-shot context poisoning; multi-turn trust escalation; indirect prompt injection; authority impersonation; all with source citations to Wei et al. 2023, Hubinger et al. 2024, Anil et al. 2024, Greshake et al. 2023, Perez & Ribeiro 2022). Eight behavioral probes per category target classes of harm (misinformation, social manipulation, privacy-violation scaffolding, policy circumvention, boundary softening, meta-compliance) — explicitly excluding Category-1 harms by design. Paired baseline vs poisoned conditions, 3 replications, 240 API calls total (~USD 20, ~75 minutes), scored by a blinded Claude-as-judge on a three-dimensional rubric (refusal, harm severity, compliance quality). Full protocol, attack templates, probe battery, scoring rubric, ethics statement, runner code, judge code, and analysis code are included. Scheduled execution: 2026-04-24. **Teaser 2 — Global Defense Feasibility Simulation v0.1 (defense side).**A 500-episode × 60-tick Monte Carlo simulation over 30 heterogeneous jurisdictions comparing three regimes (no coordination; partial coalition; full Pharos architecture) under endogenous adoption dynamics with bounded-rational peer influence on a Watts-Strogatz network. All random seeds are deposited. Headline paired-counterfactual result under v0.1 parameters: S2 (Full Pharos) reduces aggregate harm vs S0 (no coordination) by −1.14 [95% bootstrap CI: −1.28, −1.00], with the coalition stabilizing at approximately 17 of 30 jurisdictions and maintaining ≥50% membership in 0.79 of post-warmup ticks. The analysis script auto-classifies the result as "Partially supported" against a pre-registered ≥0.80 stability bar; parameters were not tuned post-hoc to clear the bar. Both teasers are minimum-viable empirical anchors, not benchmarks. Each identifies a concrete extension path into a standalone journal paper (expanded probe battery + human gold-standard + open-weight comparison for the threat-side paper; game-theoretic equilibrium + historical multilateral-regime calibration for the defense-side paper). The teasers' purpose is to ground the Pharos proposal in reproducible measurements rather than in assertions. --- ## File manifest ```AIACW_Empirical_Teasers_v0.1.zip├── README_DEPOSIT.md (this file, adapted)│├── Model_Poisoning_Teaser_v0.1/│ ├── README.md│ ├── PROTOCOL.md│ ├── ATTACKS.md│ ├── PROBES.md│ ├── RUBRIC.md│ ├── ETHICS.md│ ├── EXECUTIVE_SUMMARY.md│ ├── attacks/ (5 sanitized attack templates)│ ├── probes.json│ ├── requirements.txt│ ├── run_experiment.py│ ├── judge.py│ ├── analyze_results.py│ └── results/ (populated post-execution)│└── Global_Defense_Sim_v0.1/ ├── README.md ├── MODEL.md ├── requirements.txt ├── sim.py ├── analyze.py └── results/run_20260421_180448/ ├── config.json (500 ep × 60 tick, seed 20260424) ├── episodes.jsonl (1500 rows: 500 ep × 3 regimes) ├── summary_table.md ├── findings.md ├── figure_adoption.png ├── figure_aggregate_harm.png └── figure_failure_modes.png``` Every parameter that drives a number in `findings.md` is in either `config.json` (Sim) or in the hardcoded constants at the top of `run_experiment.py` / `judge.py` / `sim.py`. Every random draw is seeded; all runs are reproducible bit-for-bit with the same Python + numpy version. ## Reproducibility statement **Model-Poisoning Teaser.** Requires an Anthropic API key with access to `claude-opus-4-7`. Full run cost ≈ USD 20, wall-clock ≈ 75 minutes. Resume-safe (`--resume`). Identical commit + identical probe/attack files + identical sampling parameters (temperature = 1.0, max_tokens = 1024) gives statistically equivalent results; individual cells vary due to model stochasticity. Judge is temperature = 0 and is deterministic at the API level. **Global Defense Sim.** Pure CPU Monte Carlo, no API required. Full run wall-clock ≈ 3–5 minutes on a modern laptop. Single master seed (20260424) deposited in `config.json`; per-episode and per-regime seeds are deterministically derived from the master seed. Re-running `python sim.py --episodes 500 --ticks 60 --output-dir results/<new>` reproduces `episodes.jsonl` byte-for-byte under the same numpy version. --- ## Pre-registration The protocol (`PROTOCOL.md` and `MODEL.md`), the interpretation framework (`README.md` §6 and `README.md` §7 respectively), and the ethics statement (`ETHICS.md`) are pre-registered by virtue of the deposit timestamp preceding teaser-1 execution (scheduled 2026-04-24). Teaser 2 was executed on 2026-04-21 using the `MODEL.md` specification frozen the same day. Any deviations during later execution will be logged in a dated `deviations.md` inside the run directory rather than retrofitted into the protocol. --- ## Ethics and scope limits (summary; full statement in each teaser's `ETHICS.md`) - No novel attacks. Every attack category cites open-access prior work.- No Category-1 harms. Probes target moderate harm classes only.- No real targets. All named entities in probes are fabricated.- Raw model outputs retained in private workspace; this deposit contains only aggregate statistics and code.- If teaser-1 execution produces unexpectedly severe content, the run halts and the incident is disclosed to Anthropic before any external amendment to this deposit.- Both teasers are defender-favorable upper bounds — results on claude-opus-4-7 do not generalize to less-aligned models, and the results of a 30-jurisdiction abstract simulation do not generalize to actual geopolitical outcomes. These are empirical anchors, not predictions. --- ## Contact Martin K. (Hangyu Mei) — qweasdzxc12345678950@outlook.comIndependent Security Researcher, Toronto, Canada.

提供机构:
Zenodo
创建时间:
2026-04-25
二维码
社区交流群
二维码
科研交流群
商业服务