CHOI0126/llava-c3-stage1
收藏官方服务:
资源简介:
llava-c3-stage1是一个为C3架构因果性QK范数注入实验重新打包的视觉语言预训练数据集,基于原始的LLaVA-Pretrain数据集构建。它包含558,000个图像-文本对,图像来自BLIP-LAION-CC-SBU集合,文本标注以JSON和JSONL格式提供,用于多模态大语言模型的投影器训练阶段。数据集主要用于视觉语言对齐任务,支持图像到文本的转换研究。
llava-c3-stage1 is a repackaged vision-language pretraining dataset for the C3 architectural-causality QK-norm injection experiment, based on the original LLaVA-Pretrain dataset. It contains 558,000 image-text pairs with images from the BLIP-LAION-CC-SBU collection and annotations in JSON and JSONL formats, intended for the projector-only training stage of multimodal large language models. The dataset is designed for vision-language alignment tasks and supports image-to-text transformation research.
提供机构:
CHOI0126


