遇见数据集

quant_eval Golden oracle fixtures

收藏
Zenodo2026-08-19 更新2026-08-20 收录
官方服务:

资源简介:

The locked evaluation fixture set: 1,600 cases across eight agent-relevant taskfamilies, each with its deterministic ground truth. This is the input everypublished run was scored against, unchanged, so any third party can reproducethe evaluation conditions exactly. Also included is fixture_version_crosswalk.csv, mapping every published run tothe fixture SHA-256 and version label it recorded. Fixture identity: the published file is a byte-exact copy of the file used inone of the runs, not a re-serialization. Across the six published runs thefixture file appears under three distinct SHA-256 values and three versionlabels. A full structural comparison of all cases shows the only differing keyis the top-level version string; all cases, all oracle expectations, and alltrace hashes are identical. The crosswalk therefore publishes a canonicalcontent hash — SHA-256 over the fixture object with the version key removed,serialized with sorted keys and compact separators — so the equivalence can bereproduced rather than taken on trust, and flags which of the six rows describesthe file published here. About the corpus: six published runs across four base models, sixmodel-precision pairs, and three quantization schemes. Mistral-Nemo-Instruct-2407at Q4_K_M, Q5_K_M, and Q8_0 against an identical F16 baseline; Qwen2.5-7B,Qwen2.5-14B-1M, and Qwen2.5-32B at Q4_K_M against their own F16 baselines. Allfour models are Apache-2.0. Limits: the fuzz family is an adaptive trajectory evaluated from identicalstarting fixtures. Its paired test compares complete case outcomes, notidentical post-divergence prompts. Fuzz prompts are described by a publishedcontract rather than reproduced byte-for-byte, because the prompt builder is notdistributed. Licence: CC BY 4.0. No model weights are redistributed. Part of the quant_evalpublic corpus.

提供机构:
Zenodo
创建时间:
2026-08-19
二维码
社区交流群
二维码
科研交流群
商业服务