The Fathom Cognitive Atlas v0.3: Pre-registered Cross-Architecture Replication of the RLHF Attractor
收藏资源简介:
SUPERSEDED -- 2026-04-11This standalone dataset record has been absorbed into Fathom Paper v15, a unified release that contains both the Fathom paper and this atlas data bundle as sibling files on a single record under the existing Fathom concept DOI.Please cite and download the unified record instead:Paper v15: 10.5281/zenodo.19504993Concept DOI (always resolves to latest): 10.5281/zenodo.19326174The file attached to this standalone record (fathom_atlas_v0.3.zip) is byte-identical to the one included in paper v15. This standalone record remains published for archival continuity only. The Fathom Cognitive Atlas is the first public cross-architecture cognitive state atlas of open-weight language models. For every model included, we run a fixed probe set of 90 prompts across 6 categories (retrieval, reasoning, refusal, creative, adversarial, hallucination) and capture the full per-token cognitive trajectory (D, entropy, logprob, top2_margin) for every (model, prompt). The measurement primitive D = cos(h(L), WU[y]) is architecturally universal: it uses only the final-layer residual stream and the unembedding matrix, which every transformer has by construction. v0.3 contents 12 open-model captures (6 base + 6 instruct) across 3 families (Gemma, Llama, Qwen) and 2 generations (Gemma-2/Gemma-3, Llama-3.2 3B/1B, Qwen2.5 3B/1.5B). 1 exploratory closed-model capture (OpenAI GPT-4o-mini via logprob API). Probe set v0.1 - 90 prompts, SHA-pinned (34a7254920bce654), hand-curated, versioned. Analysis scripts - atlas_capture.py, atlas_analyze.py, atlas_attractor.py, atlas_attractor_n5.py, atlas_entropy.py, attractor_bootstrap.py, attractor_bootstrap_entropy.py, validate_entropy_estimator.py. All reproducible from committed JSON in under a minute on CPU. Headline result (pre-registered) The RLHF Attractor direction prediction test was pre-registered on 2026-04-10 (sealed in git before any v0.3 data was captured; see PREREG_v0.3_attractor_replication.md). The sealed decision rule: H1 SUPPORTED if and only if all three conditions hold: mean entropy early-window LOO cosine >= 0.40 permutation test p (one-sided, >= 2000 perms) < 0.05 bootstrap 95% CI lower bound > 0 At n = 6 model family pairs: mean entropy early-window LOO cosine = +0.769 permutation p = 0.0315 bootstrap 95% CI = [+0.571, +0.869] H1 SUPPORTED. All 6 of 6 families show positive LOO cosine in the pre-registered early window. The analysis script was run without modification after the decision rule was applied. Full per-family numbers, bootstrap distributions, permutation tests, and limitations are in FINDINGS_v0.3.md inside the bundle. Bulletproof audit Prior to v0.3, a rigor audit (FINDINGS_bulletproof_audit.md) identified real over-framing in the v0.2.1 writeup: the attractor claim at n=3 was structurally below significance (permutation floor p_min = 0.125). The audit added bootstrap CIs, permutation tests, and top-k entropy estimator validation across 3 model families (mean shape-correlation r = 0.902 between full-distribution entropy and top-5 API-exposed entropy). v0.3 addresses the audit's n=3 limitation by expanding to n=6 under the sealed pre-registration. Findings documents included FINDINGS_v0.1.md - seed release, 3 base models FINDINGS_v0.2.md - 6-model RLHF convergence observation FINDINGS_v0.2.1_attractor.md - first attractor formalization (superseded by v0.3) FINDINGS_entropy_bridge.md - top-k entropy proxy for closed-model extension FINDINGS_bulletproof_audit.md - rigor audit that caught v0.2.1 over-framing PREREG_v0.3_attractor_replication.md - sealed pre-registration for v0.3 FINDINGS_v0.3.md - H1 SUPPORTED, the current headline result Honest limitations n = 6 includes scale/generation variants within existing atlas families; true independent-family replication (Phi, OLMo, Mistral) is deferred to v0.4. The D-axis permutation test does not reach conventional significance even at n = 6 (p = 0.21); the primary claim is on entropy early-window only. n = 6 is small for claims about RLHF as a universal cognitive target; stable cluster-level claims require n >= 15. Companion paper The Fathom paper (Zenodo concept DOI 10.5281/zenodo.19326174) defines the measurement primitives (K, C, C_delta, S, D) and presents their pilot-scale validation. This atlas is the companion dataset that applies them at cross-architecture scale on a fixed probe set. License CC-BY-4.0 on the dataset. MIT on the analysis code. Related provisional patents: US 64/020,489 . 64/021,113 . 64/026,964.



