quant_eval Throughput telemetry
收藏资源简介:
27,370 records across 31 fields. One record per generation call across all sixpublished runs, pooled into a single union schema: token counts, timing, thetoken-count source per call, and the run and case each call belongs to. Every record carries all 31 keys. Fields absent from a source run are written asexplicit null rather than omitted, so the file is rectangular. The companionfield max_new_tokens_recorded distinguishes "the harness did not record thisfield in this run" from "the field was recorded as empty", which a bare nullcannot express — the telemetry schema gained fields mid-series, and that historyis published as data rather than left as folklore. This dataset is an independent telemetry stream. It is not derived from thescored per-case rows and is not recomputable from them; it is traceable throughthe source-bundle digests included with this record. About the corpus: six published runs across four base models, sixmodel-precision pairs, and three quantization schemes. Mistral-Nemo-Instruct-2407at Q4_K_M, Q5_K_M, and Q8_0 against an identical F16 baseline; Qwen2.5-7B,Qwen2.5-14B-1M, and Qwen2.5-32B at Q4_K_M against their own F16 baselines. Allfour models are Apache-2.0. Limits: timing is observed harness wall time on the recorded hardware andbackends. It is not a controlled throughput benchmark and not a general claimabout quantization performance at any precision on any hardware. Licence: CC BY 4.0. No model weights are redistributed. Part of the quant_evalpublic corpus.



