protoLabsAI/lab-benchmarks
收藏资源简介:
该数据集是protoLabsAI的实验室基准测试数据,用于支持模型卡中的所有数值结果。它包括发布门控结果(如量化与bf16基线的比较、配对任务集、异常值重试)、速度测试v2机制矩阵(采用InferenceMAX风格:使用种子随机数据集、客户端TTFT/TPOT的p50/p99指标、良好吞吐量)、解码深度阶梯以及一致性探测判定。方法论强调:单流数据仅在负载数据下发布;每个工作负载报告规范解码接受率(随机数据基准测试会低估约2.5倍);基于LLM判断的测试套件需通过确定性测试门控。测试工具为protoLabsAI/protoLab评估框架。每行数据中的vram_gb表示发布的权重文件大小(GB),作为质量与VRAM图表的可追溯X轴;GGUF行使用特定量化变体的文件(如NVFP4 gguf,而非同一仓库中的Q8_0)。数据集遵循CC-BY-4.0许可证,鼓励引用、使用和讨论。
This dataset comprises lab benchmarks for protoLabsAI, providing traceable data for all numbers on model cards. It includes release-gate results (e.g., quantization vs. bf16 baseline, paired task sets, outliers retrialed), speed-test-v2 regime matrices (InferenceMAX-style: seeded random dataset, client-side TTFT/TPOT p50/p99, goodput), decode-at-depth ladders, and coherence-probe verdicts. Methodology: single-stream-only numbers are published only with load numbers; spec-decode acceptance rate is reported per workload (random-data benches understate it by ~2.5x); LLM-judged suites are gated behind deterministic ones. Harness: protoLabsAI/protoLab evals. The vram_gb per row indicates the published weight-file size in GB (the downloadable artifact), serving as the traceable X-axis for quality-vs-VRAM charts; GGUF rows use specific quant variant files (e.g., NVFP4 gguf, not Q8_0 from the same repo). Licensed under CC-BY-4.0 for citation, reuse, and discussion.




