遇见数据集

EcoCompute: LLM Inference Energy on NVIDIA RTX 4090 (Ada) — FP16/NF4/INT8, 0.5B–7B, direct NVML, with paired perplexity, INT8 repeatability and a two-card NF4 re-test

收藏
Zenodo2026-09-25 更新2026-10-01 收录
官方服务:

资源简介:

Direct NVML energy measurements of decode-phase LLM inference on NVIDIA GeForce RTX 4090 cards (Ada, 24 GB), for five open models from 0.5B to 7B parameters in FP16, NF4 and INT8, with a measured run-to-run coefficient of variation and a two-physical-card re-test of the 3B NF4 anchor. Version 4 adds a fourth session (2026-09-24/25): the Qwen2.5-3B NF4 anchor re-tested three times on two physical RTX 4090 cards, on the current software stack (torch 2.14.0, CUDA 13.0, bitsandbytes 0.50.2 — deliberately the same versions as the RTX 5090 sister record 10.5281/zenodo.22855133). The rental instance migrated hosts across a restart, turning a planned same-card repeatability run into a two-card study. The July anchor stands. The three trials read +3.5%, +0.7% and +0.5% (mean +1.6%) against the July single-trial +0.8% — across a card change and a full stack change. NF4 at 3B on Ada is break-even-to-slightly-worse; the 3B anchor is unchanged by the current stack. The -15.1% reading of the 2026-08-19 session was not reproduced. None of the three current-stack trials comes within 15 points of it, on either card. Its software stack and experimental conditions differ from the present run and no specific error was identified, so it is retained as a divergent historical observation and excluded from current-stack conclusions; the mechanism of that session's offsets (its INT8 readings are likewise a factor of two above July's) remains open. No direction flip at the 3B anchor under the current stack. On the same stack, the 5090 reads this cell at -8.0/-7.4% while the 4090 reads +0.5 to +3.5%. Only the 3B anchor was re-measured; whether the full Ada crossover curve moves under this stack awaits the 1.5B and 7B points — the first controlled same-stack cross-card comparison in this series. Card and session gaps, described — not decomposed. Same card, fresh session: 0.2 pp. Different cards, same protocol: 2.8 pp — while the two cards' FP16 baselines agree within 1%. Card 1 was measured only once, so this is not a formal variance decomposition, and for card-population questions the physical sample is n_card = 2. It still motivates requiring multiple physical cards, not just multiple sessions, per architecture before calling a micro-benchmark stable. Whole-run power traces are archived for the first time in this record family: schema 1.3 sidecars (10 Hz, from before model load, with phase markers) for all six arms. Re-integrating each over its declared measurement window reproduces the reported NVML-counter energies within 0.15%, and any other window can be re-cut post hoc. Retained from version 3: the INT8 repeatability session (2026-08-20; CV of the FP16-normalised delta 0.6-3.9%, so the 100-140-point cross-session gap is 30-50x the run-to-run noise), the erratum on v1's derived vs_fp16_pct column (no measurement changed), and the 2026-08-19 session pairing each of ten measured configurations with a same-run teacher-forcing perplexity for both the quantized model and its FP16 baseline, scored on a fixed vendored public-domain text (SHA-256 22ac091a6383740d30f8e41ae144032c873c772db5e1c901112c1c330fdc5504) after the power sampler stopped. The two axes disagree: INT8 costs almost no perplexity (+0.5% to +1.2%, except Qwen2.5-3B at +3.8%) but 106% to 595% more energy, while NF4 saves energy at larger sizes yet damages the language model (+9.5% to +27.6% perplexity). They are reported side by side and deliberately not combined into a single "quality-adjusted energy" index. Scope and limitations. Every value is a real hardware measurement (basis = measured, measurement_source = direct-nvml). GPU-package power only, sampled by NVML at 10 Hz and integrated over the decode loop: this is not wall-plug power and excludes CPU, DRAM, PSU losses, PUE and CO2e. n = 1 per configuration in the 2026-07-24 and 2026-08-19 sessions; n = 3 for the five INT8 configurations of 2026-08-20 and for the single NF4 3B cell of 2026-09-24/25 (across two physical cards) — still small samples and well short of the 10 repetitions the main protocol asks for. The 10 decode iterations inside a single run are integrated into one energy total and are not independent trials. Batch size 1, 256 generated tokens, greedy decoding. The perplexity columns replicate bit-for-bit because teacher forcing is deterministic: the quality axis has CV = 0 by construction and is not independently replicated. Perplexity is a proxy for language-model damage, not a downstream-task quality guarantee, and its absolute value is specific to this text and to each model's own tokenizer. The sessions must not be pooled; what reproduces across them is the shape (the NF4 penalty falls monotonically with model size and crosses over into savings; INT8 never saves energy), not the magnitudes. Software stacks differ across sessions (A: torch 2.4.1/bnb 0.45.5; B/C: torch 2.5.1/bnb 0.43.3; D: torch 2.14.0/bnb 0.50.2) and every session's stack is recorded per row. Known metadata gaps, disclosed: the CPU block of environment.txt in the 2026-08-20 raw archive is empty; the driver version of the 2026-07-24 session appears only in that session's raw run_metadata; and in the new session the trial-B power-state log was captured post hoc on the same card (clocks and power caps are card-level defaults in every log). Measured with the open EcoCompute container (https://github.com/hongping-zh/ecocompute-mlcube), report schemas ecocompute-energy/1.1 (sessions A-C) and 1.3 with power-trace sidecars (session D). Interactive curves: https://quantenergy.tech. Supplements, and does not replace, the main EcoCompute dataset (10.5281/zenodo.19647290). Not a certified MLPerf/MLCommons benchmark; energy is reported in an MLCommons-style format, and MLCOMMONS, MLPERF and MLCUBE are trademarks of MLCommons Association, referenced nominatively.

提供机构:
Zenodo
创建时间:
2026-09-25
二维码
社区交流群
二维码
科研交流群
商业服务