遇见数据集

FilmBench

收藏
魔搭社区2026-08-09 更新2026-08-09 收录
官方服务:

资源简介:

# FilmBench — Video Generation Benchmark Dataset > ## 📢 Update (2026-08-02) > > - **Added English prompt files**: `filmbench_prompts_en.csv` is now available with English prompts. > - `filmbench_prompts_en.csv` (1,169 rows): Prompt-level table with English prompts. Columns: `uid`, `task`, `movie_type` (English), `en_prompt`, `reference_url`. > > ## 📢 Update (2026-07-31) > > - **Fixed a batch of misaligned prompts** in `filmbench_videos.csv`: the `zh_prompt` column has been recalibrated against the manually verified prompt source. > - **Added a new file `filmbench_prompts_zh.csv`**: a deduplicated prompt-level table (1,169 rows, one per prompt) for users who only need the prompts without the per-model video manifest. Columns: > > | Column | Description | > |--------|-------------| > | `uid` | Unique prompt ID, consistent with `filmbench_videos.csv`, e.g., `t2v-1`, `r2v-405` | > | `task` | Task type: `t2v` or `r2v` | > | `movie_type` | Cinematic genre (Chinese) | > | `movie_name` | Film title reference (Chinese) | > | `reference_type` | R2V reference type: `场景` (scene), `角色` (character), or `道具` (prop). Empty for T2V. | > | `zh_prompt` | Chinese prompt text | > | `image_path` | Relative path to the reference image (R2V only), e.g., `r2v_reference_image/r2v-405.png` | > > - The English-prompt columns have been temporarily removed from `filmbench_videos.csv`. **The English version of the prompts will be updated before 2026-08-02 (AOE).** FilmBench is a professional film-grade benchmark for evaluating video generation models. It covers two tasks: **T2V** (text-to-video, 515 prompts) and **R2V** (reference-to-video, 654 prompts with reference images), spanning 20 cinematic genres across Chinese Films and International Films markets. ## 1. Evaluation Framework FilmBench uses a three-level evaluation taxonomy co-designed with faculty from the Beijing Film Academy: ![FilmBench Evaluation Taxonomy](assets/newdim.png) | Level | Count | Description | |-------|-------|-------------| | **L1** (Axes) | 3 | Instruction Following, Temporal Continuity, Aesthetic Quality | | **L2** (Components) | 12 | Shot, Character & Performance, Scene, Spatial Continuity, Temporal Continuity, Character Continuity, Audio Continuity, Basic Quality, Shot Editing, Performance, Audio Quality, Visual Following (R2V only) | | **L3** (Sub-metrics) | 30–38 | Fine-grained scoring dimensions (e.g., shot scale, camera movement, composition, character appearance, scene fidelity, etc.) | Scores are aggregated bottom-up: **L3 → L2 → L1 → Overall** (arithmetic mean at each level). All scores are on a 0–100 scale. ## 2. Dataset Statistics | Item | Count | |------|-------| | Total prompts | **1,169** (T2V: 515, R2V: 654) | | Total videos | **9,213** | | T2V videos | 515 prompts × 9 models = 4,635 | | R2V videos | 654 prompts × 7 models = 4,578 | | R2V reference images | **654** (one per R2V prompt) | | Cinematic genres | 20 | | Models | 10 | ### Model Name ↔ Folder Mapping | Display Name | Folder Name | Tasks | |-------------|-------------|-------| | Seedance 2.0 | `Seedance2-0` | T2V, R2V | | HappyHorse 1.1 | `HappyHorse1-1` | T2V, R2V | | HappyHorse 1.0 | `HappyHorse1-0` | T2V, R2V | | Kling 3.0 | `KlingV3` | T2V only | | Kling 3.0 Omni | `KlingV3Omni` | T2V, R2V | | Vidu Q3 Pro | `ViduQ3Pro` | T2V only | | Veo 3.1 | `Veo3-1` | T2V, R2V | | Grok Imagine Video | `GrokImagineVideo` | T2V, R2V | | Vidu Q2 Pro | `ViduQ2Pro` | R2V only | | Hailuo 2.3 | `Hailuo2-3` | T2V only | ## 3. Overall Rankings Rankings are computed as the average `machine_overall` score across all prompts for each model. ![Overall Ranking (T2V & R2V)](assets/overall_ranking.png) ### Overall (All Models) | Rank | Model | Avg Score | Tasks | Videos | |------|-------|-----------|-------|--------| | 1 | Seedance 2.0 | 87.66 | T2V, R2V | 1,169 | | 2 | HappyHorse 1.1 | 86.35 | T2V, R2V | 1,169 | | 3 | HappyHorse 1.0 | 85.82 | T2V, R2V | 1,169 | | 4 | Kling 3.0 | 85.80 | T2V | 515 | | 5 | Kling 3.0 Omni | 82.90 | T2V, R2V | 1,169 | | 6 | Vidu Q3 Pro | 80.94 | T2V | 515 | | 7 | Veo 3.1 | 79.38 | T2V, R2V | 1,169 | | 8 | Grok Imagine Video | 78.67 | T2V, R2V | 1,169 | | 9 | Vidu Q2 Pro | 71.45 | R2V | 654 | | 10 | Hailuo 2.3 | 68.94 | T2V | 515 | ## 4. Directory Structure ``` filmbench/ ├── video/ # Generated videos, organized by model │ └── <ModelFolder>/ │ └── <ModelFolder>_<uid>.mp4 ├── r2v_reference_image/ # Reference images for R2V task │ └── <uid>.png ├── filmbench_videos.csv # Video manifest (Chinese prompts) ├── filmbench_prompts_zh.csv # Prompt-level table (Chinese only) ├── filmbench_prompts_en.csv # Prompt-level table (English only) ├── machine_scores_en.jsonl # Per-video machine scoring results (English fields) ├── machine_scores.jsonl # Per-video machine scoring results (Chinese fields) └── README.md ``` **Naming conventions:** - **uid**: `<task>-<row_id>`, e.g., `r2v-405`, `t2v-1` - **Video filename**: `<ModelFolder>_<uid>.mp4`, e.g., `Seedance2-0_r2v-405.mp4` - **Reference image**: `<uid>.png`, e.g., `r2v-405.png` (R2V only) ## 5. filmbench_videos.csv Columns | Column | Description | |--------|-------------| | `uid` | Unique prompt ID, e.g., `r2v-405`, `t2v-1` | | `task` | Task type: `t2v` or `r2v` | | `model` | Model display name, e.g., `Seedance 2.0` | | `movie_type` | Cinematic genre (Chinese), e.g., `华语剧情`, `国际动作` | | `movie_name` | Film title reference (Chinese) | | `reference_type` | R2V reference type: `场景` (scene), `角色` (character), or `道具` (prop). Empty for T2V. | | `zh_prompt` | Chinese prompt text | | `video_path` | Relative path to the video file, e.g., `video/Seedance2-0/Seedance2-0_t2v-1.mp4` | | `image_path` | Relative path to the reference image (R2V only), e.g., `r2v_reference_image/r2v-405.png` | ## 6. machine_scores_en.jsonl Fields Each line is a JSON object representing one video's machine scoring results across all evaluation dimensions. All field names are in English. | Field | Type | Description | |-------|------|-------------| | `video_filename` | string | Video filename, e.g., `Seedance2-0_t2v-1.mp4` | | `uid` | string | Prompt ID, e.g., `t2v-1` | | `task` | string | `t2v` or `r2v` | | `model` | string | Model display name, e.g., `Seedance 2.0` | | `movie_type` | string | Cinematic genre, e.g., `Chinese Drama`, `International Action` | | `reference_type` | string | R2V reference type: `scene`, `character`, or `prop` (empty for T2V) | | `machine_overall` | float | Overall score (0–100), averaged across all L1 axes | | `machine_L1_dims` | dict | L1 axis scores. Keys: `Instruction Following`, `Temporal Continuity`, `Aesthetic Quality` | | `machine_L2_dims` | dict | L2 component scores. Keys: `Shot`, `Character & Performance`, `Scene`, `Spatial Continuity`, `Temporal Continuity`, `Character Continuity`, `Audio Continuity`, `Basic Quality`, `Shot Editing`, `Performance`, `Audio Quality`, `Visual Following` (R2V only), `Audio` (R2V only) | | `machine_L3_dims` | dict | L3 sub-metric scores (30–38 keys depending on task). Key format: `<L3_name>-<L2_name>-<L1_name>`, e.g., `Shot Scale-Shot-Instruction Following` | Example: ```json { "video_filename": "Seedance2-0_t2v-1.mp4", "uid": "t2v-1", "task": "t2v", "model": "Seedance 2.0", "movie_type": "Chinese Drama", "reference_type": "", "machine_overall": 85.50, "machine_L1_dims": { "Instruction Following": 98.44, "Temporal Continuity": 85.71, "Aesthetic Quality": 72.35 }, "machine_L2_dims": { "Shot": 96.88, "Character & Performance": 100.0, "Scene": 100.0, "Spatial Continuity": 83.33, "Temporal Continuity": 100.0, "Character Continuity": 50.0, "Audio Continuity": 100.0, "Basic Quality": 84.72, "Shot Editing": 37.5, "Performance": 56.25, "Audio Quality": 100.0 }, "machine_L3_dims": { "Foreground-Background-Scene-Instruction Following": 100.0, "Scene-Scene-Instruction Following": 100.0, "Shot Scale-Shot-Instruction Following": 100.0, "...": "..." } } ``` ## 7. Links 📑 Paper: https://arxiv.org/abs/2607.24241v1 💻 GitHub: https://github.com/Neo-yk/FilmOps ## 8. Citation If you find this benchmark useful, please cite our paper: ```bibtex @misc{wang2026filmbenchfilmgradebenchmarkcinematic, title={FilmBench: A Film-Grade Benchmark for Cinematic Video Generation}, author={Shengyi Wang and Niantong Li and Guangzheng Hu and Hong Qi and Fei Ding and Weixu Qiao and Jinlin Wang and Xiaotong Lv and Peng Han and Zimeng Li and Fanshu Ding and Yushu Wang and Han Wu and Jingjing Chen and Chongxiao Wang and Yanhao Wu and Chenglong Huang and Xiaoqian Zhu and Jie Tian and Hua Li and Jingjing Fan and Mingshuang Tang and Zhong Li and Hengxia Qiang and Weibin Chen and Jinyang Zhen and Bing Zhao and Lin Qu and Jing Li and Hu Wei}, year={2026}, eprint={2607.24241}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2607.24241}, } ```

提供机构:
maas
创建时间:
2026-07-30
二维码
社区交流群
二维码
科研交流群
商业服务