HALLUCINATION BENCHMARK IN MEDICAL VISUAL QUESTION ANSWERING
收藏资源简介:
HALLUCINATION BENCHMARK IN MEDICAL VISUAL QUESTION ANSWERING是由伦敦大学学院创建的一个用于评估医学视觉问答模型幻觉现象的基准数据集。该数据集包含2359条数据,通过修改PMC-VQA、PathVQA和VQA-RAD三个公开数据集构建,涵盖了假问题、无答案选项和图像交换三种测试场景。数据集的创建旨在深入分析当前模型的局限性,并评估其在医学领域的应用效果,特别是减少幻觉现象,以提高临床决策支持的准确性。
The HALLUCINATION BENCHMARK IN MEDICAL VISUAL QUESTION ANSWERING is a benchmark dataset developed by University College London (UCL) for evaluating hallucination phenomena in medical visual question answering models. It is constructed by modifying three publicly available datasets: PMC-VQA, PathVQA, and VQA-RAD, and contains 2359 instances covering three test scenarios: fake questions, no answer options, and image swapping. The dataset is designed to conduct in-depth analysis of the limitations of current models, evaluate their performance in medical applications, and specifically mitigate hallucination phenomena to improve the accuracy of clinical decision support.




