遇见数据集

quant_eval Run provenance

收藏
Zenodo2026-08-19 更新2026-08-20 收录
官方服务:

资源简介:

6 rows by 41 columns. One row per published run: model identity, contractidentifiers, fixture hash and version label, decoding conditions, licence andcommercial-use status, and the SHA-256 and byte size of both weight artifacts —the full-weight baseline and the quantized variant. Also included is calibration_lineage.csv: 14 internal calibration runs recordedagainst the published run each informed. Calibration runs are never published;only their identifiers are disclosed, so the record is complete withoutreleasing provisional numbers. This makes references to those runs in theharness reports resolvable rather than dangling. The fields here are run-level metadata carried through from the source bundlesunchanged, not statistics derived from per-case rows. They are traceable throughthe source-bundle digests included with this record. Recomputable aggregateslive in the paired degradation statistics (D5) and family pass rates (D6)datasets, both derived from the per-case results dataset (D1). About the corpus: six published runs across four base models, sixmodel-precision pairs, and three quantization schemes. Mistral-Nemo-Instruct-2407at Q4_K_M, Q5_K_M, and Q8_0 against an identical F16 baseline; Qwen2.5-7B,Qwen2.5-14B-1M, and Qwen2.5-32B at Q4_K_M against their own F16 baselines. Allfour models are Apache-2.0. Limits: upstream_revision is populated for four of six runs. Where empty,upstream_revision_status records the cause. This is a documented absence, not adropped measurement — weight identity for those runs is still establishedexactly, by the artifact SHA-256 recorded in this file, but is not tied to anamed upstream commit. Decoding conditions are not uniform across models;temperature follows each publisher's own model card, and the two Modal runsrecord seed status "unsupported". Licence: CC BY 4.0. No model weights are redistributed. Part of the quant_evalpublic corpus.

提供机构:
Zenodo
创建时间:
2026-08-19
二维码
社区交流群
二维码
科研交流群
商业服务