llama.cpp Throughput for Qwen3.8-Flash-Next and gpt-oss-120b on a 62 GiB RAM Laptop — Reproducibility Artifact
收藏资源简介:
Mixed-license reproducibility artifact for a deployment-specific llama.cpp throughput benchmark on a 62 GiB RAM laptop. With the tested artifacts, engine checkout, and placement, the 67.564 GiB Qwen3.8-Flash-Next deployment produced 9.4–14.7 decode tokens/s across six observations, while the 59.034 GiB gpt-oss-120b deployment produced 8.7–9.3 decode tokens/s across three. The ordinary mmap path was unmodified and used no expert-cache optimization. The larger file's successful run is supporting context, not a new mmap-feasibility result; the report makes no state-of-the-art, causal storage, quality-preservation, or portable-performance claim. The bundle contains the report, sanitized timing and item-level measurement fields, deterministic checks, pinned model and environment provenance, frozen benchmark subsets, and full-rerun instructions. Model response prose and weights are excluded. Original paper, documentation, figures, and measurements are CC BY 4.0; authored code is MIT; benchmark and model material retains upstream terms. AI assistance: Claude (Anthropic) assisted with evaluation-harness development and experiment orchestration. Claude and Codex (OpenAI) assisted with evidence auditing, analysis and figure code, citation checking, and manuscript drafting and revision. Matthew Schwartz directed the research and is responsible for its methods, results, interpretation, claims, citations, rights, and released artifacts.



