Self-Drafting Qwen3.8-27B at Two Bits: Speed, Energy, and Memory on a 12 GB Laptop GPU — Reproducibility Artifact
收藏资源简介:
Mixed-license reproducibility artifact for Self-Drafting Qwen3.8-27B at Two Bits: Speed, Energy, and Memory on a 12 GB Laptop GPU. On the tested laptop, MTP raises EXL3 completion throughput from about 29 to 48.7-50.4 tokens/s on code/math and 35.7 on prose (E7). Sustained GPU-board efficiency is 0.559 tokens/J for EXL3 and 0.577 for GGUF (E25), without an established ranking. Reducing four allocated slots to one saves 2,176 MiB (E28); denser checkpoints reduce one edited-prompt latency from 10.1 to 5.0 s (E43). Drafting shows no measured accuracy decrease on 1,000 paired greedy items (E47), not equivalence. This is a frozen deployment benchmark with configuration changes, not a new method or current optimum. One loaded laptop, one 80 W cap, named artifacts and pinned engines. Most quality checks use 20-200 items at temperature 1; larger greedy comparisons remain finite-sample observations, not equivalence tests. Two-bit versus higher-bit accuracy is tested on 200 GSM8K items only. Most EXL3 measurements used mixed CUDA libraries; limited clean controls do not replace a full rerun. GPU-board energy excludes the host; newer engines and artifacts are untested. The bundle includes sanitized frozen measurements, source, configurations, patches, claim checks and full-rerun instructions. It excludes model weights, benchmark text and model-response text. Original paper, documentation and measurements are CC BY 4.0; authored code is MIT; upstream patch context retains its terms. AI assistance: Claude Code (Anthropic) operated experiments under the author's standing rules and pre-registrations; wrote engine patches, measurement tools and analysis scripts; drafted the experiment record, manuscript, claim ledger and bibliography; and prepared the artifact exporter and claim checker. Codex (OpenAI) reviewed the evidence and prose, checked primary-source literature and upstream status, revised the manuscript and companions, and checked the reading builds and artifact instructions. Matthew Schwartz directed the research and is responsible for its methods, results, interpretation, claims, citations, rights, and released artifacts.



