SEED-Bench-R1
收藏资源简介:
SEED-Bench-R1是一个用于系统评估多模态大语言模型(MLLMs)在视频理解任务中的后训练方法的基准数据集。它包括复杂的现实世界视频和日常规划任务,以多项选择题的形式呈现,要求模型具备高级的感知和推理能力。数据集通过三个层次的验证集(同分布、跨环境和跨环境任务)来评估模型的泛化能力,并配备了大规模的训练数据集,其中包含易于验证的真实答案。
SEED-Bench-R1 is a benchmark dataset for systematically evaluating post-training methods of multimodal large language models (MLLMs) in video understanding tasks. It includes complex real-world videos and daily planning tasks presented in the form of multiple-choice questions, which require models to possess advanced perceptual and reasoning capabilities. The dataset evaluates the generalization ability of models through three levels of validation sets: in-distribution, cross-environment, and cross-task scenarios. It is also equipped with a large-scale training dataset containing easily verifiable ground-truth answers.
SEED-Bench-R1 数据集概述
简介
SEED-Bench-R1 是一个用于评估多模态大语言模型(MLLMs)在视频理解任务中后训练方法的基准测试。该数据集专注于需要感知和逻辑推理的复杂任务,通过多层次评估框架(包括同分布、跨环境和跨环境-任务场景)来验证模型的泛化能力。
数据集内容
- 数据来源:基于 EgoPlan-Bench 和 EgoPlan-Bench2 的训练和验证数据。
- 数据类型:
- 大规模训练集
- 三级验证集:
- Level-1:同分布评估
- Level-2:跨环境评估(OOD)
- Level-3:跨环境-任务评估(OOD)
- 数据格式:多项选择题形式,包含四个候选答案。
- 主要指标:准确率(Accuracy)
训练与评估
- 基础模型:Qwen2-VL-Instruct-7B
- 训练方法:
- 强化学习(RL)
- 监督微调(SFT)
- 评估结果:RL 在数据效率和性能上优于 SFT,尤其在视觉感知方面表现突出。
数据获取
- 下载地址:HuggingFace
- 视频来源:Epic-Kitchens 和 Ego4D,使用时需遵守相关许可协议。
引用
bibtex @article{chen2025seedbenchr1, title={Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1}, author={Chen, Yi and Ge, Yuying and Wang, Rui and Ge, Yixiao and Qiu, Lu and Shan, Ying and Liu, Xihui}, journal={arXiv preprint arXiv:2503.24376}, year={2025} }
许可证
- 视频样本版权归 Epic-Kitchens 和 Ego4D 所有,使用时需遵守其许可协议。




