遇见数据集

ClerkBench v1: data, scorer and evidence capsule

收藏
Zenodo2026-09-30 更新2026-10-01 收录
官方服务:

资源简介:

Data and code release for ClerkBench: Measuring the Consistency Horizon of Language Models on Long Deterministic Workloads (Allcock, 2026). ClerkBench measures how long a chain of simple, fully specified clerical computation (billing, payroll, ledger replay, gateway spend tracking) a language model can execute byte-perfectly on every one of six attempts, with no tools. The release contains the byte-exact scorer, every scored task instance, all 5,313 raw model attempts behind the fourteen-configuration board, the aggregate data behind every published number, the pre-registration documents, and evidence-capsule/rebuild.py, which re-scores every attempt with the public scorer and reproduces the board exactly (Python standard library only, no network).

提供机构:
Zenodo
创建时间:
2026-09-29
二维码
社区交流群
二维码
科研交流群
商业服务