ClerkBench v1: data, scorer and evidence capsule
收藏资源简介:
Data and code release for ClerkBench: Measuring the Consistency Horizon of Language Models on Long Deterministic Workloads (Allcock, 2026). ClerkBench measures how long a chain of simple, fully specified clerical computation (billing, payroll, ledger replay, gateway spend tracking) a language model can execute byte-perfectly on every one of six attempts, with no tools. The release contains the byte-exact scorer, every scored task instance, all 5,313 raw model attempts behind the fourteen-configuration board, the aggregate data behind every published number, the pre-registration documents, and evidence-capsule/rebuild.py, which re-scores every attempt with the public scorer and reproduces the board exactly (Python standard library only, no network).



