遇见数据集

CHOI0126/llava-c3-stage1

收藏
Hugging Face2026-05-20 更新2026-05-31 收录
官方服务:

资源简介:

llava-c3-stage1是一个为C3架构因果性QK范数注入实验重新打包的视觉语言预训练数据集,基于原始的LLaVA-Pretrain数据集构建。它包含558,000个图像-文本对,图像来自BLIP-LAION-CC-SBU集合,文本标注以JSON和JSONL格式提供,用于多模态大语言模型的投影器训练阶段。数据集主要用于视觉语言对齐任务,支持图像到文本的转换研究。

llava-c3-stage1 is a repackaged vision-language pretraining dataset for the C3 architectural-causality QK-norm injection experiment, based on the original LLaVA-Pretrain dataset. It contains 558,000 image-text pairs with images from the BLIP-LAION-CC-SBU collection and annotations in JSON and JSONL formats, intended for the projector-only training stage of multimodal large language models. The dataset is designed for vision-language alignment tasks and supports image-to-text transformation research.

提供机构:
CHOI0126
二维码
社区交流群
二维码
科研交流群
商业服务