De-identified item-level scoring data: LLM as a second rater for standardized patient history-taking examinations
收藏资源简介:
Item-level scoring dataset for the manuscript 'Transcript-based large language model scoring as a low-cost second rater for standardized patient history-taking examinations: a retrospective adjudication study' (JMIR Medical Education, under review; manuscript ID 110683). 8,405 checklist items from 240 real high-stakes standardized-patient history-taking encounters (92 students, six cases), each scored by the on-site examiner, by three LLMs (DeepSeek-V4-Pro, DeepSeek-V4-Flash-0731, Qwen3-Next-80B-A3B-Instruct), and, for the 882 discordant items, by three blinded faculty experts. Encounter and student identifiers are anonymized; no names, audio, or transcripts are included. run_analysis.py reproduces the manuscript's headline analyses; v2 additionally includes an anonymized examiner identifier (E01-E12) and the cross-classified sensitivity analyses (two-way cluster bootstrap over students and examiners, and crossed random-effects models) requested during peer review. Authors: Jun Feng, Shaoting Wang, and Xiaoxing Gao contributed equally. Corresponding author: Xuefeng Sun (sunxfer@sina.com). Ethics approval: Peking Union Medical College Hospital Ethics Committee (No. I-26PJ0511).



