CareMedEval
收藏资源简介:
CareMedEval数据集是一个专门用于评估语言模型在生物医学领域进行批判性评估和推理任务的能力的数据集。该数据集来源于法国医学生的真实考试,包含基于37篇科学文章的534个问题。与现有基准不同,CareMedEval明确评估基于科学论文的批判性阅读和推理。在各种上下文条件下对最先进的通用和生物医学专业语言模型进行基准测试表明,这项任务的难度:开放和商业模型无法超过0.5的精确匹配率,尽管生成中间推理标记可以显著提高结果。然而,模型在关于研究局限性和统计分析的问题上仍然面临挑战。CareMedEval为基于情境的推理提供了一个具有挑战性的基准,揭示了当前语言模型的局限性,并为未来开发自动化支持批判性评估的推理技术铺平了道路。
The CareMedEval dataset is a specialized benchmark developed to assess the capabilities of language models in executing critical evaluation and reasoning tasks within the biomedical domain. Derived from real examinations for French medical students, this dataset contains 534 questions based on 37 scientific articles. Distinct from existing benchmarks, CareMedEval explicitly evaluates critical reading and reasoning grounded in peer-reviewed scientific papers. Benchmarking state-of-the-art general-purpose and biomedical-specialized language models across diverse contextual conditions demonstrates the difficulty of this task: both open-access and commercial models fail to exceed an exact match rate of 0.5, though generating intermediate reasoning tokens can substantially enhance their performance. Nonetheless, models still face challenges in addressing questions concerning research limitations and statistical analysis. CareMedEval offers a challenging benchmark for context-based reasoning, reveals the limitations of current language models, and paves the way for the future development of automated reasoning technologies that support critical evaluation.

- 1CareMedEval dataset: Evaluating Critical Appraisal and Reasoning in the Biomedical Field法国洛林大学,法国格勒诺布尔-阿尔卑斯大学,法国艾克斯-马赛大学 · 2025年



