VideoChat3-Academic2M
收藏资源简介:
VideoChat3-Academic2M是一个用于VideoChat3模型的学术视频指令数据集。该数据集通过对公开学术视频数据集进行重新标注构建,专门针对视频描述、视频问答和细粒度动作理解任务。它采用证据基础的标注增强流程,将原始数据集中的简短回答、选项标签和简洁描述重写为更丰富的指令跟随响应,这些响应明确提及视频中的可见对象、动作、场景、时间线索和支撑证据,并通过一致性过滤阶段确保重写后的标注与原始学术标签保持一致。数据集以JSONL格式提供标注文件,不包含原始视频文件,用户需要根据提供的原始视频数据集路径自行获取视频。数据整合了来自LLaVA-Video、Spoken-MIT、Vript、StarQA、Sports-QA和Perception-Test六个学术视频数据集的标注资源,并通过acadmic2M.json文件提供标注与视频源的映射关系。
VideoChat3-Academic2M is an academic video instruction dataset for the VideoChat3 model. It is constructed by re-annotating publicly available academic video datasets, specifically targeting video description, video question answering, and fine-grained action understanding tasks. The dataset employs an evidence-based annotation enhancement process that rewrites short answers, option labels, and concise descriptions from the original datasets into richer instruction-following responses. These responses explicitly mention visible objects, actions, scenes, temporal cues, and supporting evidence in the videos, and a consistency filtering stage ensures that the rewritten annotations align with the original academic labels. The dataset provides annotation files in JSONL format, does not include original video files, and users need to obtain the videos themselves based on provided paths to the original video datasets. It integrates annotation resources from six academic video datasets: LLaVA-Video, Spoken-MIT, Vript, StarQA, Sports-QA, and Perception-Test, and offers a mapping between annotations and video sources via the acadmic2M.json file.
数据集概述
- 名称: VideoChat3-Academic2M
- 许可证: Apache-2.0
- 任务类别: 视频-文本到文本
- 语言: 英语
- 用途: 该数据集是VideoChat3使用的学术视频指令数据,为视频描述、视频问答和细粒度运动理解等任务,对公开学术视频数据集进行了重新标注。
数据构建方法
采用基于证据的标注增强流程:
- 改写优化:将原本的简短答案、仅选项标签和简洁描述,改写为更丰富的指令跟随式回复,回复中包含可见物体、动作、场景、时间线索和支持性证据。
- 一致性过滤:保留与原始学术标签对齐的改写后标注。
数据内容与来源
数据集提供JSONL格式的标注文件,不包含原始视频。用户需根据以下来源自行获取视频:
| 源数据集 | 原始视频路径 |
|---|---|
| LLaVA-Video | https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K |
| Spoken-MIT | https://arxiv.org/pdf/2105.04489 |
| Vript | https://huggingface.co/datasets/Mutonix/Vript |
| StarQA | https://github.com/csbobby/STAR_Benchmark |
| Sports-QA | https://huggingface.co/datasets/HopLeeTop/Sports-QA |
| Perception-Test | https://github.com/google-deepmind/perception_test |
此外,acadmic2M.json文件提供了数据集标注与视频源之间的映射关系,可用于组织数据集结构。
引用说明
若使用该数据,请引用VideoChat3及其标注所基于的原始视频数据集。




