ERVQA
收藏资源简介:
ERVQA数据集由印度理工学院卡拉格普尔的研究团队创建,专注于医院环境中的视觉问答任务。该数据集包含4355个<图像,问题,答案>三元组,涵盖了急诊室中的多种场景。数据集的图像来源于真实的医院环境,问题和答案由医学专家手动标注,确保了数据的高质量和专业性。创建过程中,研究团队采用了半自动和手动结合的标注方法,并通过GPT-4V进行数据增强。ERVQA数据集主要用于评估大型视觉语言模型在医疗环境中的表现,旨在解决医疗人员短缺问题,提升智能医疗助手的性能。
The ERVQA dataset was created by a research team from the Indian Institute of Technology Kharagpur, focusing on visual question answering (VQA) tasks in hospital environments. This dataset contains 4355 <image, question, answer> triplets covering various scenarios in emergency departments. The images in the dataset are sourced from real hospital settings, while the questions and answers were manually annotated by medical experts to ensure high data quality and professional rigor. During the dataset creation process, the research team adopted a hybrid semi-automatic and manual annotation approach, and conducted data augmentation using GPT-4V. The ERVQA dataset is primarily used to evaluate the performance of large vision-language models in medical environments, aiming to address the shortage of medical personnel and improve the performance of intelligent medical assistants.




