OpenCoF-17K
收藏资源简介:
OpenCoF-17K是由字节跳动种子与香港中文大学联合构建的大规模视频推理数据集,旨在为链式帧推理提供多样化的时序监督。该数据集包含17,312个视频样本,覆盖国际象棋、数独、二维几何、点连线等11类任务家族,视频统一规范为480p分辨率、15帧率及81帧时长,平均提示词长度为53.2个单词,数据通过实例渲染、专家引导渲染、程序化场景合成和外部视频重构四类管道融合生成。其创建过程整合了结构化资产转换、人类专家标注、大语言模型概念生成以及物理引擎仿真等多模态技术,构建了可扩展的标准化视频推理格式。该数据集主要应用于增强视频生成模型的时空推理能力,旨在解决动态视觉逻辑建模、长时序连贯性及物理一致性等核心问题,推动面向推理任务的视频生成技术发展。
OpenCoF-17K is a large-scale video reasoning dataset jointly constructed by ByteDance Seed and The Chinese University of Hong Kong, which aims to provide diverse temporal supervision for chained frame reasoning. This dataset contains 17,312 video samples covering 11 task families including chess, Sudoku, 2D geometry, point-to-line connection and other related tasks. All videos are uniformly standardized to 480p resolution, 15 frames per second (fps) and a duration of 81 frames, with an average prompt length of 53.2 words. The data is generated through four integrated pipelines: instance rendering, expert-guided rendering, procedural scene synthesis and external video reconstruction. Its creation process integrates multimodal technologies such as structured asset transformation, human expert annotation, large language model (LLM) concept generation and physics engine simulation, establishing a scalable and standardized video reasoning format. This dataset is mainly applied to enhance the spatial-temporal reasoning capabilities of video generation models, aiming to solve core issues including dynamic visual logic modeling, long-term temporal coherence and physical consistency, and promote the development of video generation technologies for reasoning tasks.
数据集概述:OpenCoF-17K
数据集名称:OpenCoF-17K
所属项目:OpenCoF(Learning to Reason Through Video Generation)
数据集规模:包含 17,312 个视频样本。
任务类型:覆盖 11 个任务族(Task Families),专注于 Chain-of-Frame (CoF) 推理,即视频生成模型通过自身生成的帧序列的时间演化进行推理,而非依赖静态图像或外部工具。
数据构成:每个样本包含:
- 一张初始条件图像(initial conditioning image)
- 一个文本提示(text prompt)
- 一个目标推理视频(target reasoning video),标准化为 480p、15 fps、81 帧。
数据构建管道(共4条互补的构建流水线):
- 基于实例与专家引导的渲染:生成精确、受规则约束的视频,用于国际象棋、数独、几何等任务。
- 程序化场景合成:利用图形引擎渲染基于物理和3G的运动。
- 现有视频复用:增加现实世界的多样性,例如具身操作(embodied manipulation)。
数据集用途:
- 用于对基础模型 Wan2.2-I2V-A14B 进行 LoRA 微调,得到 Wan-CoF 模型。
- 进一步用于训练两种推理令牌(Reasoning Tokens)变体:Wan-CoFvt(视觉推理令牌)和 Wan-CoFtt(文本推理令牌)。
基准测试结果: 在四个外部视频推理基准(MME-CoF、Gen-ViRe、VIPER、RULER-Bench)上,微调后的 Wan-CoF 及其推理令牌变体在所有主要指标上均优于基线模型。





