quant_eval Family pass rates
收藏资源简介:
96 rows by 16 columns. One row per published run per runner per task family:case count, pass count, pass rate, a Wilson score confidence interval, and theexact gating signal conjunction used to compute the rate. Every row is recomputable in full from the per-case results dataset (D1). Thegating_signals column publishes the exact conjunction of scored signals a casemust satisfy to count as a pass, so a reader can reproduce any rate here withoutguessing at the definition — which is the difference between a published rateand a checkable one. Rates are reported separately for the full-weight and quantized runner of eachrun. Paired comparisons between them, with significance testing, are publishedin the paired degradation statistics dataset (D5). About the corpus: six published runs across four base models, sixmodel-precision pairs, and three quantization schemes. Mistral-Nemo-Instruct-2407at Q4_K_M, Q5_K_M, and Q8_0 against an identical F16 baseline; Qwen2.5-7B,Qwen2.5-14B-1M, and Qwen2.5-32B at Q4_K_M against their own F16 baselines. Allfour models are Apache-2.0. Limits: absolute pass rates are not comparable across models, because decodingtemperature follows each publisher's own model card rather than a single fixedvalue. Rates within a run are directly comparable, which is what the paireddesign requires. Diagnostic bucket scores are not gating pass rates and cannotsupport a fidelity or comparative claim. Licence: CC BY 4.0. No model weights are redistributed. Part of the quant_evalpublic corpus.



