VGIF-Bench
收藏资源简介:
VGIF-Bench是一个用于文本到视频生成的诊断性基准测试,专注于评估模型在时空指令跟随方面的能力。其核心目标是通过将每个生成提示表示为显式的时空依赖图(ST-DAG),使视频生成中常见的失败(如遗漏触发动作、错误转换对象或破坏时间顺序)变得可测量。数据集包含223个提示,覆盖8个宏观领域和38个微观领域,例如产品展示、电影叙事、物理交互、具身表现和现实世界场景。每个数据样本以JSONL格式提供,包含长格式生成提示和领域元数据、类型化的时空依赖图节点和边、带有布尔依赖表达式的原子问答对、用于评估电影摄影、视觉纯净度、运动流畅性和物理遵循性的提示特定自动评估维度,以及对生成视频进行诊断解释的指导。该基准仅包含测试集,用于评估规范,不包含训练标签。数据集旨在支持对视频生成模型在复杂时空推理和指令跟随方面的细粒度、可解释性评估。
VGIF-Bench is a diagnostic benchmark for text-to-video generation, focusing on evaluating models capabilities in spatiotemporal instruction following. Its core objective is to make common failures in video generation (such as missing triggered actions, incorrect object transformations, or violating temporal order) measurable by representing each generation prompt as an explicit spatiotemporal dependency graph (ST-DAG). The dataset contains 223 prompts, covering 8 macro domains and 38 micro domains, such as product showcases, movie narratives, physical interactions, embodied performances, and real-world scenes. Each data sample is provided in JSONL format and includes long-form generation prompts and domain metadata, typed nodes and edges of the spatiotemporal dependency graph, atomic question-answer pairs with Boolean dependency expressions, prompt-specific automatic evaluation dimensions for assessing cinematography, visual purity, motion fluency, and physical adherence, as well as guidance for diagnostic interpretation of generated videos. The benchmark only includes a test set for evaluation specifications and does not contain training labels. The dataset is designed to support fine-grained, interpretable evaluation of video generation models in complex spatiotemporal reasoning and instruction following.




