遇见数据集

LLaVA-CC3M-Pretrain-595K

收藏
OpenCSG2024-07-19 更新2026-01-19 收录
官方服务:

资源简介:

LLaVA Visual Instruct CC3M Pretrain 595K数据集是CC-3M数据集的一个子集,它通过更平衡的概念覆盖分布进行过滤,并包含BLIP合成字幕作为参考。该数据集主要用于视觉指令调整中的特征对齐预训练阶段,旨在构建具备GPT-4视觉/语言能力的大型多模态模型。数据包括图像与字幕对构建的多模态对话,以及图像索引、文件名、URL、原始CC-3M字幕和合成BLIP字幕等元数据。数据集的使用必须遵守CC-3M和BLIP的许可协议。主要用于大型多模态模型和聊天机器人的研究,目标用户是计算机视觉、自然语言处理、机器学习和人工智能领域的研究人员和爱好者。

The LLaVA Visual Instruct CC3M Pretrain 595K dataset is a subset of the CC-3M dataset. It is filtered to achieve a more balanced concept coverage distribution and includes BLIP-generated synthetic captions as references. This dataset is primarily utilized for the feature alignment pre-training stage in visual instruction tuning, with the goal of constructing large multimodal models that possess GPT-4-level vision and language capabilities. The data consists of multimodal dialogues built from image-caption pairs, alongside metadata such as image indices, filenames, URLs, original CC-3M captions, and synthetic BLIP captions. Usage of this dataset must adhere to the license agreements of CC-3M and BLIP. It is mainly applied to research on large multimodal models and chatbots, targeting researchers and enthusiasts in the fields of computer vision, natural language processing, machine learning, and artificial intelligence.

提供机构:
AIWizards
创建时间:
2024-07-19
搜集汇总
数据集介绍
LLaVA-CC3M-Pretrain-595K 数据集图片
背景与挑战
背景概述
LLaVA-CC3M-Pretrain-595K是CC-3M数据集的子集,通过更平衡的概念覆盖分布进行过滤,并包含BLIP合成字幕作为参考。该数据集用于视觉指令调整中的特征对齐预训练,旨在开发具备GPT-4视觉/语言能力的大型多模态模型,主要服务于计算机视觉和自然语言处理等领域的研究人员。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务