When Retrieval Hurts — Benchmark, Corrected Grader, and Evaluation Harness for On-Premise Open-Weight LLM Campus Energy Analytics (v2)
收藏资源简介:
Reproducibility artifacts for the paper "When Retrieval Hurts: Leakage-Controlled Evaluation of On-Premise Open-Weight LLMs for Affordable Campus Energy Analytics" (IEEE Transactions on Industrial Informatics, 2026). This version adds the corrected grader (ground-truth self-consistency 215/215), a template-disjoint (skeleton-cluster) holdout split with zero residual leakage, the four-condition decomposition harness (base / +output-contract / +RAG / +self-correction), the per-query latency and energy harness, and the raw model outputs behind the paper's headline result. The 215-question, 10-category benchmark is grounded in one year of measured campus-microgrid electricity data with Taiwan Power Company time-of-use tariffs and an academic-calendar load model. The raw metering data are operationally sensitive and are NOT redistributed; the harness regenerates all ground-truth values from the schema. Code: MIT; benchmark/data: CC BY 4.0.



