ImageNet-R
收藏资源简介:
该数据集基于ImageNet-R扩展,旨在应用于多模态增量学习任务。通过查询多模态大语言模型(如InstructBLIP)生成图像的描述,将图像分类数据集转化为多模态数据集。数据集的创建过程涉及对图像和文本的联合处理,以确保在有限的内存缓冲区中存储更多有代表性的样本。该数据集主要用于解决多模态增量学习中的灾难性遗忘问题,通过高效的样本存储和知识回放,提升模型的鲁棒性和效率。
This dataset is an extension of ImageNet-R, targeting multimodal incremental learning tasks. It converts image classification datasets into multimodal datasets by querying multimodal large language models (e.g., InstructBLIP) to generate image captions. The dataset creation process involves joint processing of images and texts, aiming to store more representative samples within limited memory buffers. This dataset is primarily designed to address the catastrophic forgetting problem in multimodal incremental learning, and enhances model robustness and efficiency via efficient sample storage and knowledge replay.

- 1Exemplar Masking for Multimodal Incremental Learning国立阳明交通大学、谷歌、Atmanity · 2024年



