MM-CamObj
收藏资源简介:
MM-CamObj数据集由上海交通大学创建,专门用于解决视觉语言模型在复杂场景,特别是伪装对象场景中的挑战。该数据集包含两个子集:CamObj-Align和CamObj-Instruct,分别用于视觉语言对齐和指令微调。CamObj-Align包含11,363个高质量的图像-文本对,旨在向模型注入丰富的伪装场景知识。CamObj-Instruct则包含11,363张图像和68,849个多样化的对话,用于增强模型在伪装场景中的指令跟随能力。数据集的创建过程包括从多个经典数据集中精心挑选图像,并利用GPT-4o生成详细的描述和对话。MM-CamObj数据集主要应用于评估和提升视觉语言模型在伪装对象识别、定位和计数等任务中的性能。
MM-CamObj is a dataset created by Shanghai Jiao Tong University, specifically tailored to address the challenges faced by vision-language models (VLMs) in complex scenarios, especially those involving camouflaged objects. The dataset comprises two subsets: CamObj-Align and CamObj-Instruct, which are designed for vision-language alignment and instruction tuning respectively. CamObj-Align contains 11,363 high-quality image-text pairs, aiming to inject rich knowledge of camouflage scenarios into models. CamObj-Instruct, by contrast, includes 11,363 images and 68,849 diverse dialogues, which is used to enhance the model's instruction-following capabilities in camouflage-related scenarios. The dataset creation process involves carefully curating images from multiple classic datasets and generating detailed descriptions and dialogues via GPT-4o. The MM-CamObj dataset is primarily utilized to evaluate and improve the performance of vision-language models in tasks such as camouflaged object recognition, localization and counting.
MM-CamObj
数据集概述
- 名称: MM-CamObj
- 全称: MM-CamObj: A Comprehensive Multimodal Dataset for Camouflaged Object Scenarios
- 来源: ARXIV24
- 描述: 这是一个用于“MM-CamObj: A Comprehensive Multimodal Dataset for Camouflaged Object Scenarios”的官方代码仓库。
数据集状态
- 发布状态: 代码和数据集即将发布。




