EcoCompute — LLM Inference Energy on NVIDIA RTX 4090 (Ada)
收藏官方服务:
资源简介:
This dataset reports the decode-phase inference energy of five open large language models (0.5B–7B parameters) on a single NVIDIA GeForce RTX 4090 (Ada Lovelace, 24 GB), across three precisions — FP16, NF4 (4-bit), and INT8 (8-bit). Energy is measured by direct on-device NVML power sampling at 10 Hz and integrated over the generation loop; every value is a real hardware measurement (basis = measured, measurement_source = direct-nvml), not an estimate. It extends the existing EcoCompute Ada data (previously an RTX 4090D anchor limited to 0.5B–3B, FP16/NF4) in two ways: INT8 on Ada (the prior Ada anchor had no INT8; INT8 was only available on A800). 7B models (the prior Ada anchor stopped at 3B).
提供机构:
Zenodo创建时间:
2026-07-24



