Physician ratings of sampled large language model responses to HIV-related patient questions
收藏资源简介:
This dataset contains the raw data and evaluation scores for the study assessing the performance of large language models in HIV patient queries. Large language models (LLMs) are increasingly used for health information. We evaluated 48 initial Turkish-language responses generated by three LLM services to 16 common HIV questions under a default-interface, unassisted zero-shot condition at one time point. Each response was independently rated by 124 infectious diseases physicians for accuracy, expert-perceived currency, comprehensiveness, and physician-perceived understandability. Favorable ratings (scores ≥4) were common for accuracy (ChatGPT 90.9%, Gemini 94.7%, Grok 88.6%) but less frequent for comprehensiveness (75.8%, 89.4%, and 77.2%, respectively); the sampled Gemini responses had the highest proportions in all four domains. Overall separation was greatest for comprehensiveness (Kendall’s W=0.480, 95% CI 0.389–0.579) and smallest for physician-perceived understandability (W=0.207, 95% CI 0.124–0.316). Directly estimated single-measures consistency ICCs were low (0.043–0.222); higher panel-average coefficients (0.847–0.973) concern aggregation across 124 physicians. These comparisons concern the fixed response sets and physician panel; model identity was fully confounded with presentation order. They do not establish performance for new questions, repeated generations, or later versions. Descriptive review of the Turkish transcripts identified clinically important omissions, incomplete qualification, terminology confusion, and limited safety-netting, supporting professional review of patient-facing HIV information. The attached data file includes the specific prompts used, the responses generated by the models, and the clinical evaluation scores.



