本README文件描述了一个用于存储和整理医疗人工智能基准测试评估结果的目录结构。该目录包含三个主要基准测试的输出:1) EHR-Bench(文本型电子健康记录基准),包含1800行数据和45个任务,支持单次(oneshot)和多轮智能体(agentic)两种运行模式;2) AgentEHR-Bench,包含600行数据和6个基于MIMIC数据集的任务,仅支持多轮智能体模式;3) MM Bench
Background Electronic Health Record (EHR) systems allow health care facilities to provide better care to patients and improve overall provider efficiency. They are vital for low-and middle-income coun
Background & objectives Screening for hepatitis C virus is the first critical decision point for preventing morbidity and mortality from HCV cirrhosis and hepatocellular carcinoma and will ultimately