MIRCaps full dataset for training MIRCaps-VLM
收藏官方服务:
资源简介:
Dataset Splits: Train 112k images, 784k global captions 973k cropped images, 2.8M region captions Val 14k images, 98k global captions 196k cropped images, 395k region captions Test 14k images, 98k global captions 182k cropped images, 302k region captions Note: This dataset contains images whose licenses do not permit redistribution. Please do not redistribute this dataset. Please refer to the following link to access the publicly available version of the dataset: - Paper: MIRCaps: A Large-Scale Mixed-Domain Dataset with Image-Level and Region-Level Captions for Fine-Grained Vision-Language Learning (https://arxiv.org/abs/2606.21419) - Download link: https://zenodo.org/records/20418601
提供机构:
Zenodo创建时间:
2026-08-04



