ShotQA
收藏资源简介:
ShotQA是一个大规模的多模态数据集,旨在促进视觉语言模型对电影语言的深入理解。该数据集由约7万对高质量的电影图像和视频片段中的问答对组成,涵盖8个核心的电影摄影维度,包括镜头大小、镜头构图、相机角度、镜头尺寸、光照类型、光照条件、构图和相机运动。数据集的构建过程包括数据收集与预处理、标注员培训、问答标注和验证。该数据集可以用于训练和评估视觉语言模型,以提高模型对电影摄影技巧的理解能力。
ShotQA is a large-scale multimodal dataset designed to advance the in-depth understanding of cinematography by vision-language models. It consists of approximately 70,000 high-quality question-answer pairs paired with movie images and video clips, covering 8 core cinematography dimensions including shot scale, shot composition, camera angle, shot size, lighting type, lighting conditions, framing, and camera movement. The dataset construction process includes data collection and preprocessing, annotator training, question-answer annotation, and validation. This dataset can be utilized to train and evaluate vision-language models, thereby enhancing their ability to comprehend cinematographic techniques.
ShotBench 数据集概述
数据集简介
- 名称: ShotBench
- 领域: 视觉语言模型(VLMs)的电影语言理解评估
- 数据规模: 超过3.5k专家标注的QA对
- 数据来源: 200+部奥斯卡提名电影的图像和视频片段
- 覆盖维度: 8个电影摄影核心维度
核心维度
- 镜头尺寸 (Shot Size)
- 镜头构图 (Framing)
- 摄像机角度 (Camera Angle)
- 镜头焦距 (Lens Size)
- 灯光类型 (Lighting Type)
- 灯光条件 (Lighting Condition)
- 画面构图 (Composition)
- 摄像机运动 (Camera Movement)
数据集特点
- 首个大规模电影摄影理解多模态数据集ShotQA(约70k高质量QA对)
- 包含24个主流VLMs的评估结果
- 提出新型VLM模型ShotVL(使用SFT和GRPO训练)
评估结果
- 最佳开源模型: Qwen2.5-VL-72B-Instruct(59.1%平均准确率)
- 最佳商业模型: GPT-4o(59.3%平均准确率)
- ShotVL-7B模型: 70.1%平均准确率(SOTA)
关键发现
- 约半数模型总体准确率低于50%
- 开源与商业模型性能差异不大
- 同系列中更大模型通常表现更好
- 强模型在所有维度表现均衡
- SFT+GRPO训练策略效果最佳
相关资源
- 论文: arXiv:2506.21356
- 代码: 未提供具体链接
- 数据集: 需从Huggingface下载完整版本

- 1ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models上海人工智能实验室 · 2025年



