SEED-Bench-R1
收藏资源简介:
SEED-Bench-R1是一个针对视频理解设计的多模态大型语言模型(MLLM)的评估基准,由香港大学和腾讯PCG ARC Lab共同创建。该数据集包含大量真实世界的日常活动视频,以及需要逻辑推理的多样化问题。SEED-Bench-R1的验证集分为三个层级,用于评估模型在不同泛化水平下的表现。数据集的问题设计要求模型理解开放式的任务目标,跟踪长期任务进展,感知复杂的实时环境状态,并利用世界知识进行推理以规划下一步行动。
SEED-Bench-R1 is an evaluation benchmark for Multimodal Large Language Models (MLLMs) designed for video understanding, co-created by The University of Hong Kong and Tencent PCG ARC Lab. This dataset includes a vast collection of real-world daily activity videos, alongside diverse questions that require logical reasoning. The validation split of SEED-Bench-R1 is divided into three hierarchical levels, which are used to evaluate model performance at different levels of generalization. The questions in the dataset are designed to require models to comprehend open-ended task objectives, track long-term task progress, perceive complex real-time environmental states, and leverage world knowledge for reasoning to plan subsequent actions.




