MedCalc-Eval
收藏资源简介:
MedCalc-Eval 是一个全面且具有挑战性的评估基准,用于评估大型语言模型在医疗计算方面的能力。该数据集包含了超过700个不同的临床计算任务,分为基于方程的计算和基于规则的评分系统两种类型,涵盖了包括内科、外科、儿科、重症监护、妇产科学、急诊医学、神经学、心脏病学、肺病学、泌尿学等多个临床专业。MedCalc-Eval 的创建旨在解决现有基准在评估大型语言模型在医疗计算方面的能力时的局限性,提供更准确、更全面的评估。该数据集的创建过程基于精确的方程和基于规则的评分系统,旨在为医学专业人员提供可靠的证据决策支持。
MedCalc-Eval is a comprehensive and challenging evaluation benchmark designed to assess the medical calculation capabilities of large language models (LLMs). This dataset contains over 700 distinct clinical calculation tasks, categorized into two types: equation-based calculations and rule-based scoring systems. It covers multiple clinical specialties including internal medicine, surgery, pediatrics, intensive care, obstetrics and gynecology, emergency medicine, neurology, cardiology, pulmonology, urology, and others. MedCalc-Eval was developed to address the limitations of existing benchmarks in evaluating the medical calculation capabilities of LLMs, providing more accurate and comprehensive assessments. The dataset is built upon precise equations and rule-based scoring systems, aiming to offer reliable evidence-based decision support for medical professionals.




