IMD-11
收藏资源简介:
IMD-11数据集是由南京理工大学泰州科技学院等机构的研究团队创建,包含1,637,795条图像-文本对。该数据集通过Llama模型生成图像描述,旨在为多模态学习提供丰富的图像-文本对数据。数据集的内容涵盖了11个公共图像数据集,数据量庞大,适用于少样本图像分类任务。数据集的创建过程包括使用Llama模型生成图像描述,并通过对比学习进行预训练。IMD-11数据集的应用领域主要集中在计算机视觉和多模态学习,旨在通过图像和文本的互补信息提升模型在少样本分类任务中的表现。
The IMD-11 dataset was constructed by a research team from institutions including Taizhou College of Nanjing University of Science and Technology, and comprises 1,637,795 image-text pairs. Image captions are generated via the Llama model for this dataset, which aims to provide abundant image-text pair data to support multimodal learning research. The dataset covers 11 public image datasets, boasts a large-scale data volume, and is suitable for few-shot image classification tasks. The development process of the IMD-11 dataset includes generating image captions using the Llama model and conducting pre-training via contrastive learning. The primary application domains of the IMD-11 dataset are computer vision and multimodal learning, with the objective of enhancing model performance on few-shot classification tasks by leveraging the complementary information between images and texts.

- 1IDEA: Image Description Enhanced CLIP-Adapter南京理工大学泰州科技学院, 西交利物浦大学, 昆山杜克大学 · 2025年



