遇见数据集

corpus-access-bench: open-book versus closed-book performance of language models across a capability range

收藏
Zenodo2026-08-21 更新2026-10-01 收录
官方服务:

资源简介:

Open-book versus closed-book performance of language models, measured across a capability range. Sixty doctoral-level questions in the history of economic thought, sat by four language models under five ways of reaching the same specialist corpus: cold with no tools at all, the open web, the corpus as PDFs with text search, a hand-built index of it (196 entries, 1613 cross-references), and the corpus cut into one file per entry. Every copy was graded blind, out of 60, by three judges from a different model family, against decision rules frozen before each run. Three results, against an instrument that resolves about 2 marks (established by an anchor copy re-graded across three rounds at 56.0, 57.8, 55.3): The stronger the model, the less corpus access matters. Best minus worst access: 20.3 for a local Qwen 3.8 27B, 4.5 / 1.9 / 1.7 for three frontier models. One frontier model returns 58.5 cold, 58.5 on the PDFs and 58.5 on the file directory. The fully equipped 27B does not replace a frontier model for this task. It scores 44.9; a frontier model with no tools at all scores 56.0, and the local run costs 6100-7200 s against 905-1250 s. The hand-built index never paid, in any model (-0.3, +1.3, -1.4, +2.2 against plain text search, all at or barely over the floor). The file-per-entry directory was the only access layer to move anything by double digits, and its sign flips by model: +7.3 for the 27B, -9.0 for another. Method. Each campaign committed its predictions and decision threshold before the draw; two predictions were falsified and are reported as falsified. After the real copies arrived, a salted copy in the same register carrying six defects sealed in advance was blinded into the pile: had the panel missed them, the panel's verdicts would have been discarded and the judges replaced, not the conclusion. Every panel passed. What is and is not here. The forty discriminating questions, the answer key and the blinding keys are sealed - the study measures what a model knows without the corpus, and that quantity dies the day the questions enter a training corpus - and are recorded by SHA-256 in the deposit so any campaign can prove which version it sat. The twenty-question control stratum is published whole: statements, answer key, every arm's answers, every judge's per-question line. The corpus studied is a standard multi-volume reference work in the history of economic thought; it is described by role, never named, and not redistributed. Code (harness, blinding, aggregation, local llama.cpp runner) is MIT; text and data are CC BY 4.0.

提供机构:
Zenodo
创建时间:
2026-08-21
二维码
社区交流群
二维码
科研交流群
商业服务