ShotBench
收藏资源简介:
ShotBench是一个用于评估视觉语言模型对电影语言理解的综合基准,包含超过3.5k专家注释的QA对,源自200多部备受好评的电影(主要是奥斯卡提名电影)的图像和视频片段,涵盖八个不同的电影摄影维度。
ShotBench is a comprehensive benchmark for evaluating the visual language understanding of language models, containing over 3.5k expert-annotated QA pairs derived from image and video clips of more than 200 highly acclaimed films (primarily Oscar-nominated films), covering eight distinct dimensions of cinematic cinematography.
ShotBench数据集概述
数据集简介
- 名称:ShotBench
- 类型:视觉语言模型(VLM)评估基准
- 领域:电影摄影语言理解
- 数据来源:200+部奥斯卡提名电影的图像和视频片段
核心内容
- 数据规模:包含超过3.5k专家标注的QA对
- 覆盖维度:8个电影摄影维度
- 镜头大小(SS)
- 镜头构图(SF)
- 摄像机角度(CA)
- 镜头尺寸(LS)
- 灯光类型(LT)
- 光照条件(LC)
- 镜头构图(SC)
- 摄像机运动(CM)
相关资源
- 扩展数据集:ShotQA-70k(约70k高质量QA对)
- 预训练模型:
- ShotVL-3B
- ShotVL-7B
- 论文:ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
评估结果
- 测试模型:24个领先的VLM(包括开源和专有模型)
- 最佳表现:
- GPT-4o:平均准确率59.3%
- ShotVL-7B:平均准确率70.1%(当前SOTA)
数据获取
- HuggingFace地址:
- ShotBench测试集:https://huggingface.co/datasets/Vchitect/ShotBench
- ShotQA-70k数据集:https://huggingface.co/datasets/Vchitect/ShotQA
引用格式
bibtex @misc{ liu2025shotbench, title={ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models}, author={Hongbo Liu and Jingwen He and Yi Jin and Dian Zheng and Yuhao Dong and Fan Zhang and Ziqi Huang and Yinan He and Yangguang Li and Weichao Chen and Yu Qiao and Wanli Ouyang and Shengjie Zhao and Ziwei Liu}, year={2025}, eprint={2506.21356}, achivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2506.21356}, }




