quant_eval Efficiency and footprint
收藏资源简介:
6 rows by 22 columns. One row per published run: stored weight artifact bytesbefore and after quantization, compression ratio, size reduction fraction,observed evaluation wall-time ratio with an explicit direction label, and tokenthroughput for both runners. Direction is published as a field rather than inferred from the value, becausenot every measured pair is a speedup. Four of the six pairs ran faster underquantization, and two ran slower. Compression, by contrast, is consistent acrossevery pair. The honest summary is that quantization reliably buys footprint andbuys latency only on some substrates. The figures here are not derived from the per-case results dataset. Storedartifact byte counts and observed wall time are recorded by the harness at runtime and carried through from the source bundles unchanged; they are traceablethrough the source-bundle digests included with this record, not recomputablefrom per-case rows. Pass-rate aggregates, which are recomputable, live in thepaired degradation statistics (D5) and family pass rates (D6) datasets. About the corpus: six published runs across four base models, sixmodel-precision pairs, and three quantization schemes. Mistral-Nemo-Instruct-2407at Q4_K_M, Q5_K_M, and Q8_0 against an identical F16 baseline; Qwen2.5-7B,Qwen2.5-14B-1M, and Qwen2.5-32B at Q4_K_M against their own F16 baselines. Allfour models are Apache-2.0. Limits: runtime figures are observed harness wall time on the recorded hardwareand backends. They are not a controlled throughput benchmark and not a generalclaim about quantization performance at any precision on any hardware. Themonetary cost ratio is suppressed in every row because per-runner cost was notrecorded; do not read the token volume proxy as a price. Compression is measuredon the stored weight artifact only and excludes runtime memory. Licence: CC BY 4.0. No model weights are redistributed. Part of the quant_evalpublic corpus.



