VideoChat3-Training-Data-Annotations
收藏资源简介:
VideoChat3-Stage3-Training-Data 是 VideoChat3 项目用于训练其视频多模态大语言模型(Video MLLM)的标注数据集。该数据集包含了从 Stage 0 到 Stage 3 全部四个训练阶段所使用的所有标注文件。数据组织分为三个主要部分:`training_data_annotations` 存储各训练阶段的标注数据;`videochat3_data_annotations` 包含来自各个源数据集的独立标注文件,其中每个数据条目均包含 `source` 字段以指明其原始出处;`training_data_source_tables` 则提供了源数据集的列表及路径信息。用户需根据提供的源数据链接自行下载对应的视频、图像等多媒体内容以进行完整训练。值得注意的是,本数据集中的 Stage 0 和 Stage 2 标注数据与团队先前发布的 VideoChat3-Academic2M 和 VideoChat3-LV116K 数据集存在部分重叠,为方便使用已重新整合于此。该数据集专为视频-文本到文本(video-text-to-text)任务设计,适用于训练和复现高效的通用视频理解模型。
VideoChat3-Stage3-Training-Data is an annotation dataset for training the video multimodal large language model (Video MLLM) in the VideoChat3 project. It includes all annotation files used across the four training stages from Stage 0 to Stage 3. The data is organized into three main parts: `training_data_annotations` stores annotation data for each training stage; `videochat3_data_annotations` contains independent annotation files from various source datasets, with each entry including a `source` field to indicate its original origin; `training_data_source_tables` provides a list and path information for the source datasets. Users need to download the corresponding video, image, and other multimedia content based on the provided source data links to complete the full training. Notably, the Stage 0 and Stage 2 annotation data in this dataset partially overlap with the previously released VideoChat3-Academic2M and VideoChat3-LV116K datasets, which have been re-integrated here for convenience. This dataset is specifically designed for video-text-to-text tasks and is suitable for training and reproducing efficient general-purpose video understanding models.
数据集概述:VideoChat3-Stage3-Training-Data
该数据集为 VideoChat3 模型在四个训练阶段(Stage 0 至 Stage 3)所使用的全部标注文件集合。
- 许可证:Apache-2.0
- 任务类型:视频-文本到文本(video-text-to-text)
- 语言:英语(en)
数据组织
数据集包含以下三个主要组成部分:
| 组件 | 说明 |
|---|---|
training_data_annotations |
各阶段的训练标注数据 |
videochat3_data_annotations |
单个数据集的标注文件 |
training_data_source_tables |
原始数据集及其路径信息 |
在 videochat3_data_annotations 中,Stage 0 和 Stage 2 的部分标注数据与此前发布的数据集 VideoChat3-Academic2M 和 VideoChat3-LV116K 存在重叠。为方便下载与使用,此处已重新上传。
数据使用说明
- 可通过 source-data links 下载训练所需的视频、图像及其他多媒体数据。
- 在
videochat3_data_annotations中,每个条目均提供source字段,指示该条目的来源数据集。 - 为便于使用收集和整理的高质量开源数据集复现 Stage 3 训练,另提供更详细的介绍页面:VideoChat3-Stage3-Training-Data
引用
如使用该数据,请引用 VideoChat3 论文及标注所使用的原始视频数据集。
相关链接




