MCG-NJU/VideoChat3-LV116k
收藏资源简介:
VideoChat3-LV116K是用于VideoChat3的长视频指令数据集,旨在通过长时上下文监督来补充短学术视频数据,其中证据可能稀疏、延迟并分布在多个视频片段中。该数据集通过长视频合成流水线构建:筛选候选长视频的视觉质量、语义内容和时序连贯性;将视频分割成可管理的时序片段,对每个片段进行注释和质量检查;将验证后的片段描述组装成全视频监督。基于此证据记录,流水线合成长视频字幕、长视频问答数据和时序标注。数据以JSONL注释文件形式提供,大多数原始视频未重复存储,但SciVideo视频已包含在仓库中,其他来源视频需从原始数据集中解析。数据来源包括CinePile、LongVideoDB、SciVideo_Long和SciVideo_Short数据集。
VideoChat3-LV116K is a long video instruction dataset for VideoChat3, which aims to supplement short academic video datasets via long-context supervision, where supporting evidence may be sparse, delayed and scattered across multiple video segments. This dataset is constructed through a long video synthesis pipeline: first, filter the visual quality, semantic content and temporal coherence of candidate long videos; then, split the videos into manageable temporal segments, and conduct annotation and quality inspection for each segment; finally, assemble the validated segment descriptions into full-video supervision signals. Leveraging this evidence recording mechanism, the pipeline generates long video subtitles, long video question-answering (QA) data and temporal annotations. The data is provided in the form of JSONL annotation files. Most raw videos are not stored redundantly, but SciVideo videos are already included in the repository, while videos from other sources need to be parsed from the original dataset. The data sources include CinePile, LongVideoDB, SciVideo_Long and SciVideo_Short datasets.




