MEDAGENTSBENCH
收藏资源简介:
MEDAGENTSBENCH是一个专门设计用于评估复杂医疗推理任务的基准,由耶鲁大学等机构的研究人员创建。该数据集从七个成熟的医疗数据集中精心挑选出复杂问题,这些问题需要多步骤的临床推理、诊断制定和治疗计划。数据集旨在解决现有评估中存在的三个关键问题:简单问题过多、抽样和评估协议不一致、缺乏性能、成本和推理时间的系统分析。数据集包含了来自MedQA、PubMedQA、MedMCQA、MedBullets、MMLU、MMLU-Pro、MedExQA和MedXpertQA等数据集的问题,通过严格的筛选过程确保问题难度,并包含了医学专业人士的人工注释来验证推理深度要求。
MEDAGENTSBENCH is a benchmark specially designed for evaluating complex medical reasoning tasks, developed by researchers from institutions including Yale University. This dataset carefully selects complex questions from seven well-established medical datasets, which require multi-step clinical reasoning, diagnostic formulation and treatment planning. The dataset aims to address three critical issues in existing evaluation benchmarks: overabundance of simple questions, inconsistent sampling and evaluation protocols, and the absence of systematic analysis of model performance, computational cost and inference time. The dataset includes questions sourced from MedQA, PubMedQA, MedMCQA, MedBullets, MMLU, MMLU-Pro, MedExQA and MedXpertQA. It adopts strict screening procedures to guarantee the difficulty of the questions, and incorporates manual annotations from medical professionals to validate the required depth of reasoning.

- 1MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning耶鲁大学,斯坦福大学,UT Southwestern医学中心 · 2025年



