OPEN-PMC-18M
收藏资源简介:
OPEN-PMC-18M是一个大规模高质量的生物医学视觉语言数据集,包含1800万个临床相关的子图-标题对,涵盖了放射学、显微学和可见光摄影。该数据集的创建过程涉及从BIOMEDICA语料库中筛选出600万个图像-标题对,然后使用基于Transformer的对象检测模型DAB-DETR从复合图像中提取子图,最终得到高质量的图像-标题对。该数据集的应用领域包括医学视觉语言模型和表示学习,旨在解决医学图像-文本对齐问题,提高模型的性能和临床实用性。
OPEN-PMC-18M is a large-scale, high-quality biomedical vision-language dataset containing 18 million clinically relevant subgraph-title pairs, covering radiology, microscopy, and visible light photography. The construction of this dataset involves first filtering 6 million image-title pairs from the BIOMEDICA corpus, then extracting subgraphs from composite images using the Transformer-based object detection model DAB-DETR, ultimately yielding high-quality image-title pairs. Application scenarios of this dataset include medical vision-language models and representation learning, aiming to address the medical image-text alignment task and improve model performance and clinical utility.




