遇见数据集

quant_eval Paired degradation statistics

收藏
Zenodo2026-08-19 更新2026-08-20 收录
官方服务:

资源简介:

48 rows by 28 columns. One row per published run per task family: the pairedpass-rate difference between the quantized and full-weight runners with a 95%confidence interval, the two-sided exact McNemar test, the complete discordancebreakdown, and a semantic-cluster-adjusted delta and interval. Every row is recomputable in full from the per-case results dataset (D1) usingthe published gate definitions. The pass counts, deltas, and the entirediscordance table follow directly from the per-case outcomes; no intermediateartifact is required. The cluster-adjusted columns are published alongside the row-level statisticsrather than in place of them. One task family contains semantic duplicates amongits fixtures, so row-level intervals and p-values may be optimistic if theduplicates are dependent. Publishing both, side by side, lets a reader see wherea row-level result does not survive clustering — a self-audit no competingevaluation currently performs. Headline result: on Mistral-Nemo-Instruct-2407, evaluated against an identicalF16 baseline on identical fixtures, the number of task families showing astatistically significant difference is six of eight at Q4_K_M, zero of eight atQ5_K_M, and zero of eight at Q8_0. The behavioral cost of quantization on thismodel is incurred almost entirely in the single step from Q5_K_M to Q4_K_M,while most of the storage saving is already banked at Q5_K_M. About the corpus: six published runs across four base models, sixmodel-precision pairs, and three quantization schemes. Mistral-Nemo-Instruct-2407at Q4_K_M, Q5_K_M, and Q8_0 against an identical F16 baseline; Qwen2.5-7B,Qwen2.5-14B-1M, and Qwen2.5-32B at Q4_K_M against their own F16 baselines. Allfour models are Apache-2.0. Limits: the Qwen scale ladder is confounded — the 7B run used local llama.cppwith a fixed seed, while the 14B-1M and 32B runs used Modal, which records seedstatus as "unsupported". Scale and substrate are not separated by this corpus.Cross-model comparison of absolute pass rates is confounded by temperature,which follows each publisher's own model card. Within-run pairing is unaffected,which is what the paired test requires. The fuzz family is an adaptivetrajectory; its paired test compares complete-case outcomes, not identicalpost-divergence prompts. Licence: CC BY 4.0. No model weights are redistributed. Part of the quant_evalpublic corpus.

提供机构:
Zenodo
创建时间:
2026-08-19
二维码
社区交流群
二维码
科研交流群
商业服务