CI-VID
收藏资源简介:
CI-VID数据集是一个包含超过34万个样本的文本视频数据集,每个样本由一系列视频片段和文本描述组成,旨在支持连贯的多场景视频序列生成。该数据集由北京人工智能研究院和北京邮电大学的研究人员创建,视频片段来自YouTube上超过4000个精心挑选的频道,经过严格的筛选以确保视频质量。CI-VID数据集通过构建连贯的文本-视频序列,为文本和视频到视频的生成模型提供了训练数据,这些模型能够生成具有平滑视觉转换和强时间一致性的故事驱动内容。
The CI-VID dataset is a text-video dataset comprising over 340,000 samples, each of which consists of a sequence of video clips and textual descriptions, and is designed to support coherent multi-scene video sequence generation. Developed by researchers from the Beijing Academy of Artificial Intelligence and Beijing University of Posts and Telecommunications, the dataset sources its video clips from more than 4,000 carefully curated YouTube channels, with strict filtering conducted to ensure high video quality. By constructing coherent text-video sequences, the CI-VID dataset provides training data for text-to-video and video-to-video generation models, which are capable of generating story-driven content featuring smooth visual transitions and strong temporal consistency.
CI-VID: 连贯交错文本视频数据集概述
📌 数据集简介
- 名称: CI-VID (Coherent Interleaved Text-Video Dataset)
- 类型: 大规模文本-视频交错数据集
- 规模: 超过340,000条交错视频片段与丰富字幕序列
- 设计目的: 支持连贯多片段视频生成(TV2V),超越传统孤立片段-字幕对(T2V)数据集
- 核心特性:
- 学习片段内内容与片段间过渡
- 促进具有强时间与视觉连贯性的故事驱动生成
📂 数据内容
- 字幕下载: https://flagchat.ks3-cn-beijing.ksyuncs.com/runway_log/all_train_samples.jsonl
- 视频下载: https://flagchat.ks3-cn-beijing.ksyuncs.com/runway_log/ymju_interleve/
- 可视化样本: 包含于
CI-VID_samples_for_visualization/目录
📊 评估体系
1. 人工评估
- 对比模型:
- 基线模型(Emu3训练)
- CI-VID微调模型
- 评估维度:
- 一致性
- 叙事性
- 事实正确性
- 流程: 3名专业标注员通过并排匿名比较
- 可视化示例: https://flagchat.ks3-cn-beijing.ksyuncs.com/TVinterleve/visual_contrast.zip
2. VLM评估
- 评估模型: Qwen2-VL-72B-Instruct
- 评估维度:
- 风格一致性
- 实体一致性
- 背景一致性
- 视角过渡连贯性
- 文本提示对齐
- 视觉合理性
- 评分标准: 0-5分制(极差到极优)
3. 相似性评估
- 评估层级:
- 全局相似性(完整序列)
- 对象级相似性
- 数据准备:
📜 引用信息
bibtex @misc{ju2025cividcoherentinterleavedtextvideo, title={CI-VID: A Coherent Interleaved Text-Video Dataset}, author={Yiming Ju and Jijin Hu and Zhengxiong Luo and Haoge Deng and hanyu Zhao and Li Du and Chengwei Wu and Donglin Hao and Xinlong Wang and Tengfei Pan}, year={2025}, eprint={2507.01938}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2507.01938}, }

- 1CI-VID: A Coherent Interleaved Text-Video Dataset北京人工智能研究院, 北京邮电大学 · 2025年



