AIML-TUDA/CycliST
收藏资源简介:
CycliST是一个用于评估视频语言模型在循环状态转换推理能力上的合成诊断基准数据集。它专注于周期性模式,如物体运动和视觉属性的周期性变化,以填补现有视频推理基准在捕捉线性或因果结构方面的不足。数据集包含14,800个全高清视频(1920×1080,32帧/秒,每个视频5秒/160帧),以及约120,000个模板生成的问题-答案对。视频通过Blender(Cycles引擎)进行物理渲染,并附带每帧的完整地面真值(位置、缩放、旋转、颜色、空间关系)。数据集分为5个难度等级(L1到L5),从单个循环物体到多个循环物体,并增加场景范围的周期性光照变化。支持的任务包括视频问答和场景理解/描述。数据集结构包括视频、场景元数据和问题文件夹,并分为训练、验证和测试集。数据生成过程是程序化的,确保物体在场景中的有效放置。评估使用LLM法官(Llama3-70B)进行答案评分。数据集为完全合成,局限性包括不捕捉真实世界场景的细微差别、使用固定频率循环等。发布在CC BY 4.0许可证下。
CycliST is a synthetic, diagnostic benchmark for evaluating Video Language Models (VLMs) on their ability to reason over cyclical state transitions, focusing on periodic patterns in object motion and visual attributes. It addresses the gap in existing video-reasoning benchmarks that largely capture linear or causal structure. The dataset includes 14,800 Full-HD videos (1920×1080, 32 fps, 5 seconds/160 frames each) and approximately 120k template-generated question-answer pairs. Videos are rendered using Blender (Cycles engine) with physically based rendering, and come with complete per-frame ground truth (positions, scale, rotation, color, spatial relations). It features 5 difficulty tiers (L1 to L5), ranging from one cyclic object to multiple cyclic objects with added scene-wide periodic light cycles. Supported tasks include Video Question Answering (VQA) and Scene Understanding/Captioning. The dataset structure comprises videos, scene metadata, and questions folders, split into train, validation, and test sets. Data generation is procedural with backtracking placement validation. Evaluation uses an LLM judge (Llama3-70B) for scoring free-form answers. The dataset is entirely synthetic, with limitations such as not capturing real-world scene nuances and using stationary frequencies. It is released under the CC BY 4.0 license.




