LLM-Generated OpenMC Shielding Inputs: Benchmark Specifications, Generated Scripts, Execution Records and Scorers
收藏资源简介:
Everything behind the preprint Executing the Output: An Execution-Scored Benchmark of LLM-Generated Neutron Shielding Inputs Across Three Vendors: five natural-language shielding specifications, the scoring rubric and frozen reference values; all 135 single-shot samples (Claude Opus 5 and Sonnet 5 in August 2026; Claude Opus 5.5 and Sonnet 5.5, GPT-6 Astra and Sol, Gemini 3.1 Pro and 3.8 Flash, and Gemma 4 31B in September 2026) and 53 one-turn repairs, each with its generated OpenMC script, full model response and metadata; execution records and physics verdicts; the scorers and the Google Colab notebooks that ran them; the benchmark's own defect log; and the paper's source, figures, and two model reviews with their triage. ARCHIVE.md describes the layout. OpenMC statepoint files are not included.



