INST-IT Dataset
收藏资源简介:
INST-IT Dataset是由复旦大学计算机科学学院创建的一个大规模多模态实例理解数据集。该数据集包含21,000个视频和51,000张图像,共计207,000个帧级注释,旨在提升大型多模态模型在实例级理解上的能力。数据集通过GPT-4o辅助的自动化注释管道生成,强调了实例级别的视觉提示。创建过程中,首先通过视觉提示突出感兴趣的实例,然后利用GPT-4o生成细粒度的多层次注释。该数据集主要应用于提升图像和视频中实例级理解的能力,旨在解决现有模型在处理特定实例细节时的不足。
The INST-IT Dataset is a large-scale multimodal instance understanding dataset developed by the School of Computer Science, Fudan University. This dataset contains 21,000 videos and 51,000 images, with a total of 207,000 frame-level annotations, and aims to enhance the instance-level understanding capabilities of large multimodal models. The dataset is generated via an automated annotation pipeline assisted by GPT-4o, with an emphasis on instance-level visual prompts. During its creation, visual prompts are first used to highlight instances of interest, followed by the generation of fine-grained, multi-level annotations via GPT-4o. This dataset is primarily applied to improve instance-level understanding of images and videos, and aims to address the shortcomings of existing models when handling specific instance details.

- 1Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning复旦大学计算机科学学院 · 2024年



