FilmBench
收藏资源简介:
FilmBench 是一个专业电影级视频生成基准数据集,旨在评估文本到视频(T2V)和参考图像到视频(R2V)生成模型在电影内容创作方面的性能。该数据集由与北京电影学院合作设计的三级评估分类法支持,涵盖指令跟随、时间连续性和美学质量三个主要维度(L1),下分12个组件(L2)和30至38个细粒度子指标(L3),所有评分均采用0-100分制。数据集包含1,169个提示,其中T2V任务515个,R2V任务654个(每个R2V提示附带一个参考图像),共生成9,213个视频样本(T2V: 4,635个,R2V: 4,578个),覆盖20种电影类型(如中国剧情片、国际动作片等)。数据涉及10个主流视频生成模型,包括Seedance 2.0、HappyHorse 1.1、Kling 3.0等。数据集主要文件包括:1) filmbench_videos.csv:视频清单文件,包含每个视频的唯一ID(uid)、任务类型(t2v/r2v)、模型名称、电影类型、中英文提示文本、视频文件路径和参考图像路径(R2V任务)等字段;2) machine_scores_en.jsonl:机器评分结果文件,为每个视频提供从L3到L1的详细评分数据,包括总体分数(machine_overall)和各维度分数;3) video/目录:按模型文件夹组织的生成视频文件(MP4格式);4) r2v_reference_image/目录:R2V任务的参考图像(PNG格式)。该数据集适用于视频生成模型的基准测试、性能评估和细粒度质量分析,尤其关注电影领域的专业要求。
FilmBench is a professional film-level video generation benchmark dataset designed to evaluate the performance of text-to-video (T2V) and reference image-to-video (R2V) generation models in film content creation. The dataset is supported by a three-level evaluation taxonomy developed in collaboration with the Beijing Film Academy, covering three main dimensions (L1): instruction following, temporal continuity, and aesthetic quality, with 12 components (L2) and 30 to 38 fine-grained sub-indicators (L3), all scored on a 0-100 scale. The dataset contains 1,169 prompts, including 515 for T2V tasks and 654 for R2V tasks (each R2V prompt comes with a reference image), generating a total of 9,213 video samples (T2V: 4,635, R2V: 4,578), covering 20 film genres (e.g., Chinese drama, international action films). The data involves 10 mainstream video generation models, including Seedance 2.0, HappyHorse 1.1, Kling 3.0, etc. The main files of the dataset include: 1) filmbench_videos.csv: a video manifest file containing unique ID (uid), task type (t2v/r2v), model name, film genre, prompt text in both Chinese and English, video file path, and reference image path (for R2V tasks) for each video; 2) machine_scores_en.jsonl: a machine scoring result file providing detailed scoring data from L3 to L1 for each video, including overall score (machine_overall) and dimension scores; 3) video/ directory: generated video files (MP4 format) organized by model folders; 4) r2v_reference_image/ directory: reference images for R2V tasks (PNG format). This dataset is suitable for benchmarking, performance evaluation, and fine-grained quality analysis of video generation models, with a particular focus on professional requirements in the film domain.
FilmBench 数据集概述
FilmBench 是一个用于评估视频生成模型的电影级基准数据集,涵盖文本到视频(T2V) 和参考图像到视频(R2V) 两种任务,覆盖 20 种电影类型,包含中文电影和国际电影。
1. 评估框架
采用与北京电影学院联合设计的三级评估体系:
| 层级 | 数量 | 描述 |
|---|---|---|
| L1(轴) | 3 | 指令遵循、时间连续性、美学质量 |
| L2(组件) | 12 | 镜头、角色与表演、场景、空间连续性、时间连续性、角色连续性、音频连续性、基本质量、镜头剪辑、表演、音频质量、视觉跟随(仅R2V) |
| L3(子指标) | 30–38 | 细粒度评分维度(如镜头景别、摄像机运动、构图、角色外观、场景保真度等) |
- 分数自下而上聚合:L3 → L2 → L1 → 总体(各级算术平均)
- 所有分数范围为 0–100
2. 数据集统计
| 项目 | 数量 |
|---|---|
| 总提示数 | 1,169(T2V: 515, R2V: 654) |
| 总视频数 | 9,213 |
| T2V 视频 | 515 提示 × 9 个模型 = 4,635 |
| R2V 视频 | 654 提示 × 7 个模型 = 4,578 |
| R2V 参考图像 | 654(每个R2V提示一张) |
| 电影类型 | 20 |
| 模型数 | 10 |
模型名称与文件夹映射
| 显示名称 | 文件夹名称 | 参与任务 |
|---|---|---|
| Seedance 2.0 | Seedance2-0 |
T2V, R2V |
| HappyHorse 1.1 | HappyHorse1-1 |
T2V, R2V |
| HappyHorse 1.0 | HappyHorse1-0 |
T2V, R2V |
| Kling 3.0 | KlingV3 |
T2V |
| Kling 3.0 Omni | KlingV3Omni |
T2V, R2V |
| Vidu Q3 Pro | ViduQ3Pro |
T2V |
| Veo 3.1 | Veo3-1 |
T2V, R2V |
| Grok Imagine Video | GrokImagineVideo |
T2V, R2V |
| Vidu Q2 Pro | ViduQ2Pro |
R2V |
| Hailuo 2.3 | Hailuo2-3 |
T2V |
3. 总体排名
排名基于所有提示的 machine_overall 分数的平均值。
总体排名(所有模型)
| 排名 | 模型 | 平均分 | 任务 | 视频数 |
|---|---|---|---|---|
| 1 | Seedance 2.0 | 87.66 | T2V, R2V | 1,169 |
| 2 | HappyHorse 1.1 | 86.35 | T2V, R2V | 1,169 |
| 3 | HappyHorse 1.0 | 85.82 | T2V, R2V | 1,169 |
| 4 | Kling 3.0 | 85.80 | T2V | 515 |
| 5 | Kling 3.0 Omni | 82.90 | T2V, R2V | 1,169 |
| 6 | Vidu Q3 Pro | 80.94 | T2V | 515 |
| 7 | Veo 3.1 | 79.38 | T2V, R2V | 1,169 |
| 8 | Grok Imagine Video | 78.67 | T2V, R2V | 1,169 |
| 9 | Vidu Q2 Pro | 71.45 | R2V | 654 |
| 10 | Hailuo 2.3 | 68.94 | T2V | 515 |
4. 目录结构
filmbench/ ├── video/ # 生成的视频,按模型组织 │ └── <模型文件夹>/ │ └── <模型文件夹>_<uid>.mp4 ├── r2v_reference_image/ # R2V任务的参考图像 │ └── <uid>.png ├── filmbench_videos.csv # 视频清单(每行对应一个模型×提示) ├── machine_scores_en.jsonl # 每个视频的机器评分结果(英文字段) └── README.md
命名规则:
uid:<task>-<row_id>,例如r2v-405,t2v-1- 视频文件名:
<模型文件夹>_<uid>.mp4,例如Seedance2-0_r2v-405.mp4 - 参考图像:
<uid>.png,例如r2v-405.png(仅R2V)
5. filmbench_videos.csv 列说明
| 列名 | 描述 |
|---|---|
uid |
唯一提示ID,例如 r2v-405, t2v-1 |
task |
任务类型:t2v 或 r2v |
model |
模型显示名称,例如 Seedance 2.0 |
movie_type |
电影类型(37种之一),例如 Chinese Drama, International Action |
movie_name |
电影名称参考(中文) |
reference_type |
R2V参考类型:scene, character, 或 prop。T2V为空。 |
zh_prompt |
中文提示文本 |
en_prompt |
英文提示文本 |
video_path |
视频文件的相对路径,例如 video/Seedance2-0/Seedance2-0_t2v-1.mp4 |
image_path |
参考图像的相对路径(仅R2V),例如 r2v_reference_image/r2v-405.png |
en_movie_type |
movie_type的英文翻译,例如 Chinese Drama, International Action |
en_movie_name |
movie_name的英文翻译,例如 Let the Bullets Fly - Episode 1 |
6. machine_scores_en.jsonl 字段说明
每行是一个JSON对象,代表一个视频在所有评估维度上的机器评分结果。所有字段名为英文。
| 字段 | 类型 | 描述 |
|---|---|---|
video_filename |
string | 视频文件名,例如 Seedance2-0_t2v-1.mp4 |
uid |
string | 提示ID,例如 t2v-1 |
task |
string | t2v 或 r2v |
model |
string | 模型显示名称,例如 Seedance 2.0 |
movie_type |
string | 电影类型,例如 Chinese Drama, International Action |
reference_type |
string | R2V参考类型:scene, character, 或 prop(T2V为空) |
machine_overall |
float | 总体分数(0–100),所有L1轴的平均值 |
machine_L1_dims |
dict | L1轴分数。键:Instruction Following, Temporal Continuity, Aesthetic Quality |
machine_L2_dims |
dict | L2组件分数。键:Shot, Character & Performance, Scene, Spatial Continuity, Temporal Continuity, Character Continuity, Audio Continuity, Basic Quality, Shot Editing, Performance, Audio Quality, Visual Following (仅R2V), Audio (仅R2V) |
machine_L3_dims |
dict | L3子指标分数(30–38个键,取决于任务)。键格式:<L3名称>-<L2名称>-<L1名称>,例如 Shot Scale-Shot-Instruction Following |
示例: json { "video_filename": "Seedance2-0_t2v-1.mp4", "uid": "t2v-1", "task": "t2v", "model": "Seedance 2.0", "movie_type": "Chinese Drama", "reference_type": "", "machine_overall": 85.50, "machine_L1_dims": { "Instruction Following": 98.44, "Temporal Continuity": 85.71, "Aesthetic Quality": 72.35 }, "machine_L2_dims": { "Shot": 96.88, "Character & Performance": 100.0, "Scene": 100.0, "Spatial Continuity": 83.33, "Temporal Continuity": 100.0, "Character Continuity": 50.0, "Audio Continuity": 100.0, "Basic Quality": 84.72, "Shot Editing": 37.5, "Performance": 56.25, "Audio Quality": 100.0 }, "machine_L3_dims": { "Foreground-Background-Scene-Instruction Following": 100.0, "Scene-Scene-Instruction Following": 100.0, "Shot Scale-Shot-Instruction Following": 100.0, "...": "..." } }
7. 相关链接
- 论文: https://arxiv.org/abs/2607.24241v1
- GitHub: https://github.com/Neo-yk/FilmOps
8. 引用
bibtex @misc{wang2026filmbenchfilmgradebenchmarkcinematic, title={FilmBench: A Film-Grade Benchmark for Cinematic Video Generation}, author={Shengyi Wang and Niantong Li and Guangzheng Hu and Hong Qi and Fei Ding and Weixu Qiao and Jinlin Wang and Xiaotong Lv and Peng Han and Zimeng Li and Fanshu Ding and Yushu Wang and Han Wu and Jingjing Chen and Chongxiao Wang and Yanhao Wu and Chenglong Huang and Xiaoqian Zhu and Jie Tian and Hua Li and Jingjing Fan and Mingshuang Tang and Zhong Li and Hengxia Qiang and Weibin Chen and Jinyang Zhen and Bing Zhao and Lin Qu and Jing Li and Hu Wei}, year={2026}, eprint={2607.24241}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2607.24241}, }




