quant_eval Per-case behavioral results
收藏资源简介:
19,200 rows by 146 columns. One row per evaluation case per runner across sixpublished runs: the raw model output, every scored signal, per-case timing,expected and observed value pairs, the oracle trace, the fuzz audit envelope,and the decoding conditions under which the row was produced. This is the substrate of the corpus. Every aggregate published in the paireddegradation statistics dataset (D5) and the family pass rates dataset (D6) isrecomputable from this file using the gate definitions in the accompanyingdata_dictionary.json. Nothing in between is required. Eight agent-relevant task families at 200 cases each, evaluated twice per run —once against a full-weight baseline and once against a quantized variant of thesame model, case for case: json, json_multistep, mcq, mixed_brief_json,stateful_followup, toolcall, toolcall_only, and fuzz. The accompanying data_dictionary.json documents all 146 fields: applicability,data type, evidence role, gate membership, the meaning of an empty value, andobserved population counts. An empty cell is not automatically a missingmeasurement — several fields are conditional on task family, and theirempty-value meaning is recorded explicitly. About the corpus: six published runs across four base models, sixmodel-precision pairs, and three quantization schemes. Mistral-Nemo-Instruct-2407at Q4_K_M, Q5_K_M, and Q8_0 against an identical F16 baseline; Qwen2.5-7B,Qwen2.5-14B-1M, and Qwen2.5-32B at Q4_K_M against their own F16 baselines. Allfour models are Apache-2.0. Statistical comparison is paired: the two-sidedexact McNemar test on per-case outcomes, with Wilson intervals on the rates. Limits: decoding conditions are not uniform across models — temperature followseach publisher's own model card, so cross-model comparison of absolute passrates is confounded. Within-run pairing is unaffected, which is what the pairedtest requires; the conditions are published per row so they can be filtered on.Runs on the Modal substrate record seed status "unsupported" and are not exactlyreproducible. Licence: CC BY 4.0. No model weights are redistributed. Part of the quant_evalpublic corpus.



