Replication materials for "Benchmarking Frontier Large Language Models on Consumer-Facing Cost Questions: Price-Range Width, Output Consistency, and a Structured-Engine Baseline"
收藏资源简介:
Anonymized data and reproduction code for the study. Includes the 40-question cross-sectional pass, the four deep repeated questions, the twenty-question repeated set for both general-purpose models, and a self-contained, standard-library-only script (score_reproduce.py) that recomputes every reported figure: price-range width medians and distribution, the over-charge marker counts, and the sample standard deviation across repeated runs. No network access or API keys are required. See README.md to run.Version 3 (2026-09-26): figure-type-aware re-scoring and figure scripts, added for the round-1 revision of the Technical Note (Intelligent Infrastructure and Construction, manuscript iic-4476502). Nothing from versions 1 and 2 is changed or removed. New files: rescore_v2.py (classifies every man-yen figure as itemized, unit rate, or prose; computes the full span [the published rule, unchanged], the span excluding unit rates, the prose span, and the headline range = the first explicit range in running text; for repeated runs, the sample SD of the full span and the lower/upper endpoint ratios and overlap of the headline ranges; validates first that its full span reproduces the published width columns row by row), make_figures.py (Figures 1 and 2 of the revised manuscript), deep4_gpt55_runs.json (per-run headline range and full span for the four deep-repetition questions), figure1/figure2 (PNG 600 dpi and PDF), README_v3.md (rules in words, how to run). Run in a folder holding the v1 CSVs and the unzipped v2 runs/ directory: python3 rescore_v2.py. Python 3.8+, standard library only (matplotlib for the figures).



