遇见数据集

Dataset for Generative artificial intelligence models in clinical infectiousdisease consultations: a cross-sectional analysis among specialists andresident trainees

收藏
Figshare2025-02-13 更新2026-04-28 收录
官方服务:

资源简介:

In this cross-sectional analysis, researchers evaluated the performance and safety of four generative artificial intelligence (GenAI) chatbots in the context of clinical infectious disease consultations. The study involved GPT-4.0, a custom GPT-4.0 chatbot (cGPT-4) optimized via retrieval-augmented generation, Gemini Pro, and Claude 2. Forty unique clinical scenarios were created from real patient consultations and systematically anonymized and categorized into relevant sections. Six clinical experts, including specialists and resident trainees, independently evaluated 160 AI-generated responses using a 5-point Likert scale across four domains: factual consistency, comprehensiveness, coherence, and potential medical harmfulness.Results demonstrated that GPT-4.0-based models achieved significantly higher composite scores compared to Gemini Pro and Claude 2, particularly in factual accuracy and comprehensiveness. However, across all models, less than two-fifths of responses were deemed “Harmless,” raising concerns about the clinical safety of deploying these systems without supervision. Specialists consistently rated the responses more favorably than resident trainees, highlighting a discrepancy in clinical judgment. Cost analysis revealed decreasing operating expenses over time, but model performance was not directly correlated with cost. The study emphasizes that despite promising advancements, current GenAI models require further refinement and human oversight before being integrated into direct clinical care. Collectively, these findings inform future clinical application.

本项横断面分析中,研究人员针对临床感染病会诊场景下的四款生成式人工智能(Generative Artificial Intelligence, GenAI)聊天机器人的性能与安全性展开了评估。本次研究纳入了GPT-4.0、一款通过检索增强生成(Retrieval-Augmented Generation, RAG)优化的定制化GPT-4.0聊天机器人(cGPT-4)、Gemini Pro以及Claude 2。研究人员从真实的患者会诊记录中构建了40个独立的临床场景,并对其进行系统化匿名化处理,随后归类至相关分类模块中。6名临床从业人员(含专科医师与住院医师规培学员)采用李克特5级量表,从事实一致性、全面性、连贯性以及潜在医疗危害性四个维度,对160份AI生成的会诊回复进行独立评分。研究结果显示,基于GPT-4.0的模型在综合评分上显著优于Gemini Pro与Claude 2,尤其在事实准确性与内容全面性维度表现突出。不过,四款模型生成的回复中,仅有不到五分之二被评定为"Harmless",这引发了学界对未加监督便部署此类系统用于临床场景的安全性担忧。专科医师对回复的评分始终高于住院医师规培学员,这凸显出不同层级临床人员的临床判断存在差异。成本分析结果显示,模型的运营成本随时间推移呈下降趋势,但模型性能与运行成本并无直接关联。本研究强调,尽管生成式人工智能技术已取得可观进展,但当前的GenAI模型在投入直接临床应用前,仍需进一步优化并辅以人工监督。综上,本研究结果可为未来此类技术的临床应用提供参考依据。

创建时间:
2025-02-13
二维码
社区交流群
二维码
科研交流群
商业服务