遇见数据集

Positive-Result Challenge of the Honesty Harness (Levels 1-2): pre-registration, sealed assignment, ledger and code

收藏
Zenodo2026-09-25 更新2026-10-01 收录
官方服务:

资源简介:

This deposit contains the complete record of the pre-registered, mechanically blinded positive-result challenge required by §6(iii) of the Honesty Harness manuscript preregistration: the frozen challenge preregistration with its §10 verdict appended below the freeze marker (`PREREGISTRO_CHALLENGE.md`, frozen-part sha256 `79644d780df7…`, marker-included convention, with sha256 sidecar and OpenTimestamps receipt); the sealed blind assignment of 600 synthetic worlds to arms (`assignment.json`, sha256 `d693da6582ee…`, with sidecar and OpenTimestamps receipt — the `os.urandom` master seed inside it was not seen by the director until unmasking); the closed verdict ledger (`verdict_ledger.json`, sha256 `dcaf788469df…`, one verdict per world under the primary gates, with the secondary gate and the 200-world sensitivity floor recorded per world); both floors, sealed before any battery verdict (`floor.json`, `floor_200.json`); the calibration file with the full δ→edge response curve and per-λ realized edges (`calibration_delta.json`); the unabridged run log of the 600-world blind battery (`run.log`, 591.7 s, no manual intervention); the unmasked aggregates (`AGGREGADOS.json/.md`); and, under `code/`, the exact challenge code at commit `d7547ab409ad` — world generator for the three arms, sealed-assignment builder, blind runner with its mechanical guard, ledger, unmasking gate and test suite (7/7). Headline result, under the frozen gates and falsifiers: the double-gate verdict procedure detected an in-family planted signal with expected net edge of twice trading costs in 19/20 blind worlds (Wilson 95% [76%, 99%]), produced 24/500 = 4.80% false positives on pure-noise worlds against a nominal α = 5% (point estimate below α; the Wilson interval [3.25%, 7.04%] contains 5%, so the statement is not conclusive at 95%), and detected an out-of-family signal in 0/20 worlds. Falsifiers F1, F2 and F3 did not fire; the realized false-positive rate exceeded its pre-registered expectation band (0.5–3%), with the diagnosis — sampling noise of a tail quantile estimated from a 20-world floor — recorded in §10 together with the sensitivity variant (2.00% under the 200-world floor). The system under test is honesty-harness-toolkit v0.1.0 (doi:10.5281/zenodo.22846067); this record is a supplement to the Honesty Harness protocol (doi:10.5281/zenodo.21838807) and to the manuscript "The Honesty Harness: expert-directed, AI-executed empirical research with frozen criteria and asymmetric verdicts". License: CC BY 4.0. Author: Ricardo Castellanos Macias (ORCID 0009-0009-8031-589X).

提供机构:
Zenodo
创建时间:
2026-09-25
二维码
社区交流群
二维码
科研交流群
商业服务