VideoChat3-LV116k
收藏资源简介:
VideoChat3-LV116K是VideoChat3使用的长视频指令数据集,旨在通过提供更长时序上下文的监督来补充短学术视频数据,特别适用于证据可能稀疏、延迟且分布在多个视频片段中的场景。数据集通过结构化流程构建:首先筛选候选长视频(基于视觉质量、语义内容和时序连贯性),然后分割为可管理的时序片段,接着对每个片段进行标注和质量检查,最后将验证通过的片段描述组装成完整的视频监督信息。基于这个流程,合成了长视频字幕、长视频问答数据以及时间定位标注。数据集以JSONL标注文件形式提供,原始视频需从源数据集获取,源数据集包括CinePile(电影总结与问答)、LongVideoDB(长视频时间线/时间定位/总结与问答)、SciVideo_Long(SciVideo总结与问答)和SciVideo_Short(SciVideo总结与问答)四个类别。适用于视频-文本到文本任务,如视频理解、视频问答、视频总结和时间定位等需要理解长时序上下文的应用场景。
VideoChat3-LV116K is the long video instruction dataset utilized by VideoChat3, aiming to supplement short academic video datasets by providing supervision over longer temporal contexts, and is particularly suited for scenarios where evidence may be sparse, delayed, and distributed across multiple video segments. The dataset is constructed via a structured data pipeline: first, candidate long videos are screened based on visual quality, semantic content and temporal coherence; then they are segmented into manageable temporal segments; subsequently, each segment undergoes annotation and quality inspection; finally, the verified segment descriptions are assembled into complete video supervision information. Following this pipeline, long video subtitles, long video question-answering data and temporal localization annotations are synthesized. The dataset is distributed in the form of JSONL annotation files, and the original videos must be obtained from the source datasets, which cover four categories: CinePile (movie summarization and question answering), LongVideoDB (long video timeline/temporal localization/summarization and question answering), SciVideo_Long (SciVideo summarization and question answering) and SciVideo_Short (SciVideo summarization and question answering). It is applicable to video-text-to-text tasks, such as video understanding, video question answering, video summarization and temporal localization, as well as other application scenarios that require understanding of long temporal contexts.
数据集概述:VideoChat3-LV116K
许可证:Apache-2.0
任务类型:视频-文本到文本
语言:英语
数据集简介
VideoChat3-LV116K 是 VideoChat3 模型使用的长视频指令数据集,旨在补充短视频学术数据集在长时域上下文(如稀疏、延迟、分布在不同视频片段中的证据)上的监督信息。
数据构建方法
该数据集通过长视频合成流水线构建,关键步骤包括:
- 候选视频筛选:根据视觉质量、语义内容和时域连贯性过滤候选长视频。
- 视频分段与标注:将视频切分为可管理的时域片段,逐段进行标注和质量检查。
- 全视频监督合成:基于验证后的片段描述,组装生成:
- 长视频描述文本
- 长视频问答数据
- 时域定位标注
数据格式
数据集以 JSONL 格式提供标注文件(lv116k.json),其中包含:
- 数据标注与原始视频源的映射关系
- 不包含原始视频文件,用户需自行从原始数据集路径获取视频
数据来源
VideoChat3-LV116K 基于以下视频数据集构建:
| 来源数据集 | 类别 | 原始视频数据集路径 |
|---|---|---|
| CinePile | 电影摘要与问答 | https://huggingface.co/datasets/tomg-group-umd/cinepile |
| LongVideoDB | 长视频时间线/时域定位/摘要与问答 | https://huggingface.co/datasets/LongVideos/LongVideoDB-373K-Videos |
| SciVideo_Long | 科学视频摘要与问答 | 未提供 |
| SciVideo_Short | 科学视频摘要与问答 | 未提供 |
引用说明
使用该数据时,需引用 VideoChat3 以及标注所基于的原始长视频数据集。




