quant_eval Public Corpus
收藏资源简介:
The corpus-level record for quant_eval: a per-case behavioral evaluation offull-weight and quantized large language models across eight agent-relevant taskfamilies, with paired statistical testing and published ground truth. Quantized model cards state a compression ratio and stop. They do not say whatthe quantized model stops being able to do. This corpus measures that directly:identical fixtures, identical cases, a paired test per task family, and everyper-case row published so the aggregates can be recomputed by anyone. This record carries the corpus overview, licence, citation metadata, theverbatim source-bundle checksums, and the build manifest — a SHA-256 digest andbyte count for every file across all seven datasets. Cite this record for thecorpus as a whole. The seven datasets each declare isPartOf against it: D1 Per-case behavioral results — 19,200 rows x 146 columnsD2 Throughput telemetry — 27,370 records x 31 fieldsD3 Golden oracle fixtures — 1,600 cases, plus the version crosswalkD4 Run provenance — 6 rows x 41 columns, plus calibration lineageD5 Paired degradation statistics — 48 rows x 28 columnsD6 Family pass rates — 96 rows x 16 columnsD7 Efficiency and footprint — 6 rows x 22 columns Scope: six published runs across four base models, six model-precision pairs,and three quantization schemes. Mistral-Nemo-Instruct-2407 was evaluated againstan identical F16 baseline at Q4_K_M, Q5_K_M, and Q8_0, forming a three-pointprecision curve. Qwen2.5-7B-Instruct, Qwen2.5-14B-Instruct-1M, andQwen2.5-32B-Instruct were each evaluated at Q4_K_M against their own F16baselines, forming a parameter-scale ladder. All four models are Apache-2.0. Verification: all 72 source-bundle file digests were verified, and all 96 familyby runner pass rates independently recomputed from the raw per-case rows beforethe corpus was built. Every published file carries a SHA-256 in the buildmanifest. Licence: CC BY 4.0. Commercial use is permitted; attribution is required. Nomodel weights are redistributed; each evaluated model remains under its ownlicence, recorded per run in the run provenance dataset (D4).



