EHRNoteQA
收藏资源简介:
EHRNoteQA数据集是由韩国科学技术院的研究团队开发,专门用于评估大型语言模型在临床环境中的表现。该数据集包含962个独特的问题,每个问题都与特定患者的电子健康记录(EHR)临床笔记相关联。EHRNoteQA的独特之处在于它是首个采用多选项问答格式的数据集,这种设计有效地评估了在自动评估背景下大型语言模型的可靠性。此外,它要求分析多个临床笔记以回答单一问题,反映了现实世界临床决策的复杂性,其中临床医生需要审查患者历史的广泛记录。通过EHRNoteQA,研究团队对各种大型语言模型进行了全面的评估,显示了其在评估医疗应用中大型语言模型的重要性和促进大型语言模型整合到医疗系统中的关键作用。
The EHRNoteQA dataset was developed by a research team from the Korea Advanced Institute of Science and Technology (KAIST) specifically to evaluate the performance of large language models (LLMs) in clinical settings. This dataset contains 962 unique questions, each linked to the clinical notes of a specific patient’s electronic health record (EHR). What makes EHRNoteQA unique is that it is the first dataset adopting the multiple-choice question answering format, a design that effectively assesses the reliability of LLMs in the context of automated evaluation. Furthermore, it requires analyzing multiple clinical notes to answer a single question, which reflects the complexity of real-world clinical decision-making, where clinicians need to review extensive records of a patient’s medical history. Through EHRNoteQA, the research team conducted comprehensive evaluations of various LLMs, underscoring the importance of evaluating LLMs for medical applications and the critical role of this dataset in facilitating the integration of LLMs into healthcare systems.




