BIOMEDICA
收藏资源简介:
BIOMEDICA是由斯坦福大学开发的一个大规模生物医学图像-文本数据集,旨在填补生物医学领域缺乏多样化、公开可访问的多模态数据集的空白。该数据集包含超过2400万条图像-文本对,源自600万篇PubMed Central开放获取的文章,涵盖了广泛的生物医学领域,如病理学、放射学、眼科学、皮肤病学等。数据集通过专家注释和丰富的元数据(如文章标题、摘要、关键词等)进行增强,支持流式处理和高效查询。BIOMEDICA的创建过程包括从PubMed Central提取数据、生成图像特征、聚类并由专家进行注释,最终通过Hugging Face平台公开发布。该数据集的应用领域广泛,旨在推动生物医学视觉-语言模型的发展,支持零样本分类、图像-文本检索等任务,为精准医疗提供数据支持。
BIOMEDICA is a large-scale biomedical image-text dataset developed by Stanford University, aiming to fill the gap caused by the lack of diverse, publicly accessible multimodal datasets in the biomedical field. This dataset contains over 24 million image-text pairs, derived from 6 million open-access articles in PubMed Central, covering a broad spectrum of biomedical disciplines including pathology, radiology, ophthalmology, dermatology and more. The dataset is enhanced with expert annotations and rich metadata such as article titles, abstracts, keywords and other relevant information, and supports streaming processing and efficient querying. The development pipeline of BIOMEDICA includes data extraction from PubMed Central, image feature generation, clustering and expert annotation, and it is finally publicly released through the Hugging Face platform. With a wide range of application scenarios, this dataset aims to promote the development of biomedical vision-language models, support tasks such as zero-shot classification and image-text retrieval, and provide data support for precision medicine.

- 1BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature斯坦福大学 · 2025年



