LLM consistency on nursing clinical judgement: raw responses, per-cell metrics, and analysis code
收藏资源简介:
Raw API responses, per-cell and per-item metrics, analysis scripts, and figure-generation code for a two-study evaluation of the response consistency of three commercial large language models — GPT-4o (OpenAI), Claude Opus 4.5 (Anthropic), and Gemini 2.5 Flash-Lite (Google) — on nursing clinical-judgement items, framed as a deployment-safety evaluation of whether output consistency can serve as a pre-deployment safety signal. Study 2 (the primary analysis in the current manuscript) is an independent, externally authored, cross-lingual external validation on 90 publicly released items of the Japanese National Nursing Examination (Ministry of Health, Labour and Welfare; 106th–115th examinations, 2017–2026). Each item was queried 30 times at each feasible sampling temperature across the three models (990 cells × 30 trials = 29,700 analysed responses). Study 1 (the initial English-language pilot, reported as supplementary) uses 30 author-written English nursing clinical-judgement scenarios across four sampling temperatures (Claude capped at T ≤ 1.0 by the Anthropic API), yielding 9,836 valid responses out of 9,900 attempted calls. Response consistency is measured as 1 − H(R)/log2(k), the normalised complement of the Shannon entropy of the empirical response distribution over the repeated trials (k = number of answer options), and is reported jointly with accuracy. The repository documents the residual of items answered unanimously yet keyed wrong (the confidently-wrong minority) and the limits of a consistency-based triage safeguard. The raw Japanese examination PDFs are public and available from the MHLW examination page cited in the manuscript. Code is released under the MIT Licence; data and figures under CC BY 4.0.



