FiVL-Instruct
收藏资源简介:
FiVL-Instruct数据集是由英特尔实验室创建的,旨在增强视觉语言模型的视觉对齐能力。该数据集基于LLaVA-1.5-mix-665K指令调优数据集,包含665,000条结构化对话,每条对话包含多个问题和答案,且94%的对话包含图像。数据集通过GPT-4o提取关键表达,并使用GroundedSAM生成精确的分割掩码,以增强视觉信息与文本的对齐。该数据集主要用于训练和评估视觉语言模型在视觉问答任务中的表现,旨在解决模型在视觉信息利用上的不足,提升模型的解释性和准确性。
The FiVL-Instruct dataset was developed by Intel Labs to enhance the visual alignment capabilities of vision-language models (VLMs). Built upon the LLaVA-1.5-mix-665K instruction-tuning dataset, it contains 665,000 structured dialogues, each with multiple question-answer pairs, and 94% of these dialogues include associated images. The dataset extracts key expressions via GPT-4o and generates precise segmentation masks using GroundedSAM, aiming to strengthen the alignment between visual information and textual content. Primarily used for training and evaluating the performance of vision-language models on Visual Question Answering (VQA) tasks, this dataset is designed to address the deficiencies of models in leveraging visual information, and to improve the interpretability and accuracy of such models.

- 1FiVL: A Framework for Improved Vision-Language Alignment英特尔实验室 · 2024年



