NCLM-GATE-PILOT-0: pilot of the OSIRIS core speaking-gate measurement (reclassified from 'NCLM-1: Core Learning Gate Test (Day 1)')
收藏资源简介:
Pilot, not confirmatory. This record establishes no claim that the core model learns. Version 1 of this record was deposited as day 1 of a pre-registered 7-day test of whether the OSIRIS core model (osiris.nclm, a byte-level transformer trained online on mentor-model replies) improves its prediction of held-out conversation. As written, that protocol could not produce a result: PASS required 30 held-out exchanges but the schedule yielded about 3. The run also used a different mentor model (llama3.2:1b) from the one pre-registered (qwen2.5:7b), its measurement functions did not exist, and its name collided with a separately locked NCLM-1 protocol, which this record does not touch. Version 2 therefore reclassifies the data as NCLM-GATE-PILOT-0, with ERRATA.md and DEVIATIONS.md. Contents: the 11 hash-chained exchanges (8 training, 3 held out) and the version-1 files byte-identical; a paired pre/post scorer that scores every held-out item with a frozen pre-training snapshot and with the trained core; a replay of the 8 training exchanges on a frozen copy of a local core (5 seeds); and a draft, unlocked confirmatory design. Descriptive pilot result: held-out bits/byte moved from 5.36–5.59 to 5.32–5.56 (mean paired difference 0.043 bits/byte, 15/15 item×seed pairs positive), unchanged with repeated template text masked. A byte unigram fit on the same 8 replies scores 4.45–4.58, better than the core before and after. Three held-out items from one day and one mentor cannot estimate the variance needed to power a confirmatory test; see analysis/effect_size_plan.md.



