RefVideo-6M
收藏资源简介:
RefVideo-6M是一个大规模、高可靠性的参考引导视频编辑数据集,由中国电信人工智能研究院、香港中文大学等机构联合构建。该数据集包含500万视频编辑样本和100万图像编辑样本,总计600万条,视频分辨率达720p且帧数介于81至129帧,覆盖26种视频编辑任务与6种图像编辑任务,并提供约600万视觉参考,涵盖边界框、圆形、风格、纹理等多元参考类型。数据集采用逆向构建策略,以无伪影的真实视频作为编辑目标,利用12个编辑专家生成输入条件,并借助大语言模型过滤失败案例,确保监督信号的可靠性。该数据集旨在解决现有编辑数据集因依赖自动生成模型而引入伪影、以及缺乏视觉参考导致可控性不足的问题,为训练高质量参考引导视频编辑模型提供坚实基础,显著提升编辑的视觉质量、可控性与参考一致性。
RefVideo-6M is a large-scale, high-reliability reference-guided video editing dataset jointly constructed by China Telecom Artificial Intelligence Research Institute, The Chinese University of Hong Kong and other institutions. This dataset contains 5 million video editing samples and 1 million image editing samples, totaling 6 million samples overall. The videos have a resolution of 720p with frame counts ranging from 81 to 129, covering 26 video editing tasks and 6 image editing tasks, and provides approximately 6 million visual references spanning diverse types such as bounding boxes, circles, artistic styles, textures and more. The dataset adopts a reverse construction strategy: taking artifact-free real videos as editing targets, employing 12 professional editing experts to generate input conditions, and filtering failed cases with large language models (LLMs) to guarantee the reliability of supervision signals. This dataset addresses the core limitations of existing editing datasets, including artifacts introduced by reliance on automated generation models and insufficient controllability caused by the lack of visual references. It provides a solid foundation for training high-quality reference-guided video editing models, and significantly improves the visual quality, controllability and reference consistency of edited content.

- 1RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing中国电信人工智能研究院; 香港中文大学; 中山大学; 复旦大学; 清华大学 · 2026年



