Reproducibility deposit — A small language model detects behavioural faithfulness gaps that frontier judges and human raters miss
收藏资源简介:
Reproducibility package for the manuscript "A small language model detects behavioural faithfulness gaps that frontier judges and human raters miss": evaluation code, evaluation sets (including the 90,314-item auditor held-out set), tokeniser, corpus manifest, measured results, rater protocols, and the hierarchical-specificity ablation. All numbers are measured; scenario content is fully synthetic. See README_DEPOSIT.md inside the archive for the folder-by-folder manifest. Version 2 (2026-07-19): record title updated to match the current manuscript title; demonstration memory items in the retention scripts replaced with a fictional persona for privacy (an early demonstration log quoting those items was removed); four-condition experiment results unchanged; one malformed row in the held-out set repaired; per-item matched-geometry evidence added; requirements.txt and CHANGELOG.md added. See CHANGELOG.md inside the archive.



