proof-bundles
收藏资源简介:
Chinese-LLaVA是一个用于视觉语言任务的中文多模态数据集,旨在支持视觉语言模型的训练和评估,特别关注中文语境下的多模态理解与生成任务。数据集包含三个主要组成部分:训练数据、评估数据和指令微调数据。训练数据包含约100,000个图像-文本对,图像来源于COCO数据集,每个图像配有详细的中文描述文本。评估数据包含约5,000个图像-文本对,同样来自COCO数据集,用于模型性能评估。指令微调数据包含约50,000个指令-响应对,专门用于训练模型遵循复杂指令的能力。数据集支持多种视觉语言任务,包括图像描述生成、视觉问答和视觉推理。数据以JSONL格式组织,每个样本包含图像路径、问题、答案、指令类型等字段。该数据集适用于训练和评估中文多模态大语言模型,促进视觉与语言理解的交叉研究。
Chinese-LLaVA is a Chinese multimodal dataset for vision-language tasks, designed to support the training and evaluation of vision-language models, with a particular focus on multimodal understanding and generation tasks in the Chinese context. The dataset consists of three main components: training data, evaluation data, and instruction tuning data. The training data includes approximately 100,000 image-text pairs, with images sourced from the COCO dataset, each accompanied by detailed Chinese descriptive text. The evaluation data includes about 5,000 image-text pairs, also from the COCO dataset, used for model performance assessment. The instruction tuning data contains around 50,000 instruction-response pairs, specifically designed to train models in following complex instructions. The dataset supports various vision-language tasks, including image captioning, visual question answering, and visual reasoning. Data is organized in JSONL format, with each sample containing fields such as image path, question, answer, and instruction type. This dataset is suitable for training and evaluating Chinese multimodal large language models, promoting cross-disciplinary research in vision and language understanding.
- 许可证: MIT




