Medical Complex Vision Question Answering Dataset (MeCoVQA)
收藏资源简介:
MeCoVQA数据集是由百度公司、中国农业大学、自动化研究所、中国科学院和北京大学联合创建的,旨在支持复杂医学影像问答和图像区域理解的多模态数据集。该数据集包含8种模态,共计31万对问答数据,涵盖了医学影像的详细信息和语义描述。创建过程中,首先将图像的分割掩码转换为结构化元数据,然后通过大型语言模型生成图像描述,最终整合生成复杂的问答数据。MeCoVQA数据集主要应用于医学领域的多模态大语言模型研究,旨在提升医学影像的像素级理解和问答能力,解决医学影像分析中的细粒度问题。
The MeCoVQA dataset is a multimodal dataset jointly developed by Baidu, China Agricultural University, Institute of Automation, Chinese Academy of Sciences, and Peking University, designed to support complex medical visual question answering and image region understanding tasks. This dataset contains 310,000 question-answer pairs across 8 modalities, covering detailed information and semantic descriptions of medical images. During the dataset curation process, image segmentation masks were first converted into structured metadata, followed by the generation of image captions via large language models (LLMs), and finally complex question-answer pairs were generated through integration. The MeCoVQA dataset is primarily applied to multimodal large language model research in the medical domain, with the objective of enhancing pixel-level understanding and question-answering capabilities of medical images, and addressing fine-grained challenges in medical image analysis.




