AQ-MedAI/MedDialogRubrics
收藏资源简介:
MedDialogRubrics 是一个用于评估医疗大型语言模型多轮咨询能力的大规模基准测试和评估框架。与静态医疗问答基准不同,它评估医生模型是否能主动收集临床重要信息、管理对话轨迹、识别紧急预警信号,并逐步完善其诊断推理。该基准包含5,200个合成的患者案例和超过60,000个细粒度的评估标准。患者案例基于疾病知识生成,未使用真实世界电子健康记录;而评估标准候选则基于循证医学知识,通过拒绝采样过滤,并由临床专家评审和优化。
MedDialogRubrics is a large-scale benchmark and evaluation framework for assessing the multi-turn consultation capabilities of medical large language models (LLMs). Unlike static medical question-answering benchmarks, it evaluates whether a doctor model can actively gather clinically important information, manage the dialogue trajectory, identify urgent warning signs, and progressively refine its diagnostic reasoning. The benchmark contains 5,200 synthetically constructed patient cases and more than 60,000 fine-grained evaluation rubrics. Patient cases are generated from disease knowledge without accessing real-world electronic health records, while rubric candidates are grounded in evidence-based medical knowledge, filtered through rejection sampling, and subsequently reviewed and refined by clinical experts.




