Towards AI Epidemiology: reproducibility bundle for the LLM-judge live run (v1.1)
收藏资源简介:
Reproducibility bundle for the minimal application reported in Section 8 of "Towards AI epidemiology: a measurement standardisation framework for prospective risk detection" (K. Tempest-Walters). An LLM judge (gemini-2.5-pro, temperature 0) scores treatment plans from GPT-3.5, GPT-4, and a rheumatology board across 19 rheumatology vignettes on three fields: risk level (grounded in the NPSA Consequence Score), policy alignment (grounded in the applicable EULAR recommendation), and evidential alignment (grounded in the evidence base within that guideline). Each plan is scored twice for a test-retest measure, with verbose and concise variants for a verbosity probe (190 scored interactions). The archive contains the pipeline code, the 190 raw judge response logs, the analysis outputs (test-retest ICC and weighted kappa, expert-versus-AI comparison, grammar records), and the frozen judge instruction. Input vignettes and the model and board plans are reused from Labinsky et al. (2024, Rheumatology International, CC BY 4.0). The EULAR guideline PDFs are third-party copyrighted documents and are not redistributed here.



