BioVGQ
收藏资源简介:
BioVGQ数据集是基于PMC-VQA数据集建立的,包含了经过筛选的77000张清洁的生物医学图像和188000个问题-答案对。该数据集通过整合多个公共数据集,并过滤掉经过显著手动处理的图像,同时利用图像和相应的说明生成问题和答案,以确保问题-答案对与图像内容的高度相关性。数据集的建立旨在解决生物医学视觉问答中的固有偏差问题,并用于训练所提出的BioD2C模型。
The BioVGQ dataset is built upon the PMC-VQA dataset, comprising 77,000 curated clean biomedical images and 188,000 question-answer pairs. This dataset is constructed by integrating multiple public datasets, filtering out images that have undergone significant manual manipulation, and generating question-answer pairs using the images and their corresponding captions to ensure high relevance between the pairs and the image content. The dataset is designed to address the inherent bias issues in biomedical visual question answering and is used for training the proposed BioD2C model.

- 1BioD2C: A Dual-level Semantic Consistency Constraint Framework for Biomedical VQA山东大学, 青岛校区, 中国 · 2025年



