KeyFrame-Compass
收藏资源简介:
KeyFrame-Compass是一个用于评估关键帧条件视频生成模型的基准数据集。给定一个有序的关键帧图像序列和一个文本提示,模型需要生成一个视频,该视频能在指定时间位置复现预设的视觉状态,并在连续锚点之间合成连贯的运动、过渡和事件。数据集包含386个精心策划的基础样本,涵盖短视频和长视频生成。每个样本以两种条件格式(多图像有序列表和包含相同序列的单一故事板网格图像)提供相同的排序关键帧序列,并为每种格式配对两种级别的文本控制(极简提示和分段特定提示)。这种设计支持在不同生成设置下对关键帧执行情况和整体视频质量进行受控评估。样本根据视频持续时间进一步组织为短视频子集(227个样本)和长视频子集(159个样本)。每个样本沿五个主要维度进行标注:输入格式、提示控制级别、视频结构(连续的一镜到底视频或多镜头视频)、关键帧数量(3、6、9或12个视觉锚点)和应用领域(日常捕捉、产品可视化或电影叙事)。任务定义根据视频结构有所不同:对于多镜头视频,每个关键帧被分配到一个特定的镜头并具有“起始”、“结束”或“代表”角色;对于一镜到底视频,每个关键帧被分配一个目标时间戳,相邻锚点界定时间片段。提示变体包括:极简提示(仅指定定义任务所必需的信息,如关键帧顺序、简要故事概要、目标视频结构和目标时长)和分段特定提示(额外描述每个关键帧的预期时间位置、主体状态及其演变、相机语言和镜头行为等)。比较这两种变体可以将弱文本指导下的生成与明确时间约束下的指令遵循区分开来。数据集文件结构包含一个全局样本索引(samples.jsonl)以及按持续时间划分的`short/`和`long/`目录,每个样本目录包含关键帧图像、故事板网格图像、提示文件和元数据文件。
KeyFrame-Compass is a benchmark dataset for evaluating keyframe-conditioned video generation models. Given an ordered sequence of keyframe images and a text prompt, models are required to generate a video that reproduces preset visual states at specified temporal positions, while synthesizing coherent motion, transitions and events between consecutive anchor points. The dataset contains 386 carefully curated base samples covering both short and long video generation. Each sample provides the same ordered keyframe sequence in two conditional formats: a multi-image ordered list and a single storyboard grid image containing the same sequence, and pairs each format with two levels of text control (minimal prompt and segment-specific prompt). This design enables controlled evaluation of keyframe execution performance and overall video quality across different generation settings. The samples are further organized into short video subset (227 samples) and long video subset (159 samples) based on video duration. Each sample is annotated along five main dimensions: input format, prompt control level, video structure (continuous single-take video or multi-shot video), number of keyframes (3, 6, 9 or 12 visual anchors), and application domain (everyday capture, product visualization or cinematic narrative). The task definition varies according to video structure: for multi-shot videos, each keyframe is assigned to a specific shot with a "start", "end" or "representative" role; for single-take videos, each keyframe is assigned a target timestamp, with adjacent anchors defining temporal segments. Prompt variants include: minimal prompts (only information necessary to define the task, such as keyframe order, brief story outline, target video structure and target duration) and segment-specific prompts (additional descriptions of the expected temporal position of each keyframe, subject state and its evolution, camera language and shot behavior, etc.). Comparing these two variants allows distinguishing generation under weak text guidance versus instruction following with explicit temporal constraints. The dataset file structure includes a global sample index (samples.jsonl) as well as `short/` and `long/` directories categorized by duration. Each sample directory contains keyframe images, storyboard grid images, prompt files and metadata files.
数据集概述
KeyFrame-Compass 是一个用于评估关键帧条件视频生成(Keyframe-conditioned Video Generation)的基准测试(Benchmark)数据集。给定有序的关键帧图像序列和文本提示,模型需生成一段视频,该视频需在指定时间位置复现给定的视觉状态,并在相邻关键帧之间合成连贯的运动、过渡和事件。
核心构成
- 样本总数:386 个精心策划的基础样本(Base Samples),覆盖短视频与长视频生成。
- 子集划分:
- 短视频子集:227 个样本。
- 长视频子集:159 个样本。
输入格式与提示变体
每个基础样本提供两种输入格式和两种提示文本控制水平,用于多维度评估:
- 输入格式(Input Formats):
- 多图像列表(Multi-image List):将关键帧作为有序的独立图像序列输入,图像存储于
selected/目录,对应提示在multi-image-list/目录。 - 故事板网格(Storyboard Grid):将关键帧按相同顺序排列成一张网格图像(
storyboard_grid.png),对应提示在storyboard-grid/目录。
- 多图像列表(Multi-image List):将关键帧作为有序的独立图像序列输入,图像存储于
- 提示变体(Prompt Variants):
- 最小化提示(Minimal):仅指定关键帧顺序、简短故事概要、目标视频结构和时长,让模型从视觉序列推断运动与事件。
- 分段特定提示(Specific):在最小化提示基础上,进一步描述关键帧的预期时间位置、主体状态演化、镜头语言、镜头间事件及叙事进程。
样本标注维度
每个样本在五个主要维度上进行标注:
- 输入格式:多图像列表或故事板网格。
- 提示控制水平:最小化或分段特定。
- 视频结构:连续单镜头(One-take)或多镜头(Multi-shot)视频。
- 关键帧数量:3、6、9 或 12 个视觉锚点。
- 应用领域:日常拍摄(Daily capture)、产品展示(Product visualization)或电影叙事(Cinematic narrative)。
任务定义
- 多镜头视频:每个关键帧被分配至特定镜头,并承担角色(如第一帧、最后一帧或代表性帧)。
- 单镜头视频:每个关键帧沿连续轨迹分配目标时间戳,相邻关键帧之间需描述主体演化、相机运动和叙事进展。
数据组织
数据集根目录下包含 samples.jsonl 全局样本索引文件,并按时长分为 short/ 与 long/ 两个目录。每个样本目录结构如下:
| 路径 | 内容描述 |
|---|---|
selected/ |
有序的独立关键帧图像(用于多图像列表输入)。 |
storyboard_grid.png |
包含相同关键帧的单一网格图像。 |
multi-image-list/ |
适配多图像列表输入的两种提示(minimal 和 specific)。 |
storyboard-grid/ |
适配单网格图像输入的两种提示(minimal 和 specific)。 |
manifest.json |
样本级元数据和时间规范。 |
其他信息
- 语言:英语。
- 任务类别:图像到视频(image-to-video)。
- 许可协议:其他(license: other)。
- 伦理考量:数据集样本经过多模态验证、人工审核和安全审查,但建议用户在使用前自行检查内容,遵守相关许可与使用政策。




