遇见数据集

Auditing Public Evaluation Specifications and Scorer Smoke-Testability: audit corpus, coding instrument, and execution record

收藏
Zenodo2026-07-27 更新2026-08-02 收录
官方服务:

资源简介:

Evidence dataset and analysis pipeline auditing the public evaluation artifacts of 26 frozen foundation-model benchmark releases for data-science tasks. Releases are laddered as Described (A0), Accessible (A1), Sufficiently documented (R2) and Smoke-testable (E3), the last established by CPU-only, zero-inference evaluator runs. The release contains: the 35-candidate screening frame and frozen corpus registry; line-addressable evidence packets; raw and aggregated judgments from three independently operated model judges; both completed human rater files with their 63 uncollapsed disagreements and the adjudicated 156-cell gold set; golden-fixture inputs with independently derived expected values; official-environment build specifications and captured output; CPU-only execution logs; minimal-repair patches; and a single-command reproduction pipeline with a SHA-256 manifest over every released file. Headline reliability result: the pre-registered inter-coder gate was missed by both the model panel and the human panel on identical evidence, so the documentation-sufficiency endpoint is released as descriptive rather than validated, with all disagreements preserved.

提供机构:
Zenodo
创建时间:
2026-07-27
二维码
社区交流群
二维码
科研交流群
商业服务