SagaBench — Season 1 and Wave 5 run records (artifact record v1.0)
收藏资源简介:
WHAT THIS IS. The complete run records behind "The Steward's Paradox" (SagaBench AB, 2026 — published, not peer-reviewed): 2,415 Season-1 runs (23 models × 7 situations × 15 replicates) and 960 preregistered Wave-5 runs (71 situations classified before any model ran, two deployed cheap-tier models). One JSON per run with the agent's full decision log, the replay triple (engine hash, situation, edicts), the nine-component outcome receipt for the agent run and the no-agent counterfactual, and the counterfactual composite that the paper's numbers are computed from. RECOMPUTE THE PAPER'S NUMBERS IN FIVE MINUTES. scripts/recompute_season1.py and scripts/recompute_classes.py (standard-library Python 3) return every published count and rate exactly: Season 1 156 / 2,415 = 6.5 % catastrophes (a run whose counterfactual composite is below −30, i.e. the agent left its situation worse than no agent acting at all), 95 % CI [3.6, 9.9], cluster-robust calibrated [2.6, 11.1]; Wave 5 KNIFE-EDGE 87 / 500 = 17.4 % [10.6, 24.8], ROBUST 13 / 200 = 6.5 %, GROWTH 7 / 200 = 3.5 % [0.5, 7.5]. "pip install sagabench" and "sagabench verify <run>.json" checks any run's score receipt. README.md explains which intervals reproduce bit-for-bit and which reproduce up to bootstrap noise, and why. PROVENANCE. The randomness-beacon rounds (drand / League of Entropy) the situations were drawn from, the locked preregistrations with SHA-256 sidecars, the Season-1 seed-derivation recipe, the Wave-5 classification and the hash-and-timestamp commitment to its selection log, the replay-audit attestations (Season 1: 2,415 / 2,415 bit-identical replays on one runtime; Wave 5: 960 / 960 on two), and the analysis provenance of the published Wave-5 intervals. MANIFEST.json lists the SHA-256 of every file and is OpenTimestamps-anchored (MANIFEST.json.ots). WHAT IS WITHHELD, AND WHY. The simulation engine, the situation generator and the hold-out seeds are not published: an open engine lets any model be trained on the test. Wave-5 records carry no model identifier, provider, seed or rationale text, because the Wave-5 preregistration commits SagaBench never to publish per-model rates for that wave; class labels, keyed situation identifiers, decision logs and receipts are included, so every class-conditional number recomputes. Season-1 records include seeds and model identifiers; the paper reports Season 1 at group level and this record does not support per-model rankings (per-model ordering does not reproduce across independent halves of the situation set; see README). The alpha-era material announced in paper version 1.0 is not in this record; version 1.1 withdraws it. FILES. season1.tar.gz (2,415 records + manifest), wave5.tar.gz (960 records + class summary + manifest), provenance.tar.gz (26 files), scripts.tar.gz, README.md, ERRATA.md, LICENSE (CC BY 4.0; the engine is not part of this record), MANIFEST.json, MANIFEST.json.ots. Questions and corrections: info@sagabench.com.



