MedHallBench
收藏资源简介:
MedHallBench是由华威大学、克兰菲尔德大学和牛津大学的研究团队开发的一个基准数据集,专门用于评估医学大语言模型(MLLMs)中的幻觉问题。该数据集通过整合专家验证的医学案例场景和现有医学数据库构建,涵盖了广泛的医学知识和临床情境。数据集的内容包括详细的医学案例、医学文献和临床报告,确保了数据的多样性和深度。创建过程中,研究人员采用了自动标注方法,如强化学习与人类反馈(RLHF),以提高数据标注的效率和准确性。MedHallBench的应用领域主要集中在医疗保健领域,旨在解决MLLMs在生成医学信息时的幻觉问题,从而提高模型在临床环境中的可靠性和安全性。
MedHallBench is a benchmark dataset developed by research teams from the University of Warwick, Cranfield University and the University of Oxford, specifically designed to evaluate hallucination issues in medical large language models (MLLMs). This dataset is constructed by integrating expert-validated medical case scenarios and existing medical databases, covering a wide range of medical knowledge and clinical contexts. The content of the dataset includes detailed medical cases, medical literature and clinical reports, ensuring the diversity and depth of the data. During its development, researchers adopted automatic annotation methods such as Reinforcement Learning from Human Feedback (RLHF) to improve the efficiency and accuracy of data annotation. The application fields of MedHallBench are mainly focused on the healthcare sector, aiming to address the hallucination problem of MLLMs when generating medical information, thereby enhancing the reliability and safety of these models in clinical settings.

- 1MedHallBench: A New Benchmark for Assessing Hallucination in Medical Large Language Models华威大学, 克兰菲尔德大学, 牛津大学 · 2025年



