LLaVA-23k
收藏资源简介:
LLaVA(Large Language and Vision Assistant)数据集由威斯康星大学麦迪逊分校、微软研究院和哥伦比亚大学联合创建,旨在推动多模态视觉与语言理解的发展。该数据集通过利用 GPT-4 生成的指令数据,构建了首个大规模的视觉指令跟随数据集,包含约 15.8 万条多模态语言 - 图像指令跟随样本,涵盖对话、详细描述和复杂推理等多种类型。数据集的创建基于广泛存在的图像 - 文本对数据,通过设计特定的提示词,引导 GPT-4 生成与视觉内容相关的指令和回答。其应用领域广泛,主要用于训练能够理解多模态指令并完成视觉任务的通用视觉助手,如视觉问答、图像编辑、导航等。LLaVA 数据集为开发和评估多模态模型提供了丰富的资源,有助于推动机器人导航、虚拟现实交互等领域的研究。
The LLaVA (Large Language and Vision Assistant) dataset was co-created by the University of Wisconsin-Madison, Microsoft Research, and Columbia University, aiming to advance the development of multimodal vision-language understanding. This dataset constructs the first large-scale visual instruction-following dataset by leveraging instruction data generated by GPT-4, containing approximately 158,000 multimodal language-image instruction-following samples covering various types such as dialogue, detailed description, and complex reasoning. The dataset is developed based on widely existing image-text pairs, and by designing specific prompts, it guides GPT-4 to generate instructions and responses related to visual content. It has a wide range of application scenarios, mainly used to train general-purpose visual assistants that can understand multimodal instructions and complete visual tasks, such as visual question answering, image editing, navigation, and so on. The LLaVA dataset provides rich resources for developing and evaluating multimodal models, which helps promote research in fields including robot navigation and virtual reality interaction.




