Camellia054/ShareGPT4V
收藏资源简介:
ShareGPT4V Captions 1.2M 是一个由GPT4-Vision驱动的多模态字幕数据集。它旨在增强大型多模态模型在预训练和监督微调阶段的模态对齐和细粒度视觉概念感知,以推动这些模型向GPT4-Vision的能力发展。数据集包含两个主要部分:一部分由GPT4-Vision生成(ShareGPT4V),另一部分由在GPT4-Vision生成数据上训练的Share-Captioner生成(ShareGPT4V-PT)。此外,还从ShareGPT4V部分中精选了数据用于监督微调阶段。该数据集于2023年11月7日收集,主要用于大型多模态模型和聊天机器人的研究,目标用户包括计算机视觉、自然语言处理、机器学习和人工智能领域的研究人员和爱好者。
ShareGPT4V Captions 1.2M is a set of GPT4-Vision-powered multi-modal captions data. It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Multi-Modal Models (LMMs) during both the pre-training and supervised fine-tuning stages. This advancement aims to bring LMMs towards GPT4-Vision capabilities. The dataset includes components generated by GPT4-Vision (ShareGPT4V) and by a Share-Captioner trained on GPT4-Vision-generated data (ShareGPT4V-PT), with a curated subset for supervised fine-tuning. It was collected on November 7, 2023, and is intended for research on large multimodal models and chatbots, targeting researchers and hobbyists in computer vision, NLP, machine learning, and AI.



