MMR-V
收藏资源简介:
MMR-V数据集是由中国科学院自动化研究所等单位创建的多模态视频深度推理基准。该数据集包含317个视频和1257个任务,涵盖了动画、电影、哲学、电视、生活和艺术六大类别。数据集的特点是要求模型进行长距离多帧推理,不仅要感知问题帧,还要分析远离问题帧的证据帧。任务包括显式推理和隐式推理两种类型,旨在评估模型在视频理解、情感识别、因果推理、序列结构推理、反直觉推理、跨模态迁移推理和视频类型与意图推理等方面的能力。数据集的创建遵循多帧、深度推理和现实性三个原则,视频来源广泛,任务设计严谨,旨在推动多模态推理能力的研究。
The MMR-V dataset is a multimodal video deep reasoning benchmark developed by the Institute of Automation, Chinese Academy of Sciences and other institutions. It consists of 317 videos and 1257 tasks, covering six major categories: animation, film, philosophy, television, life, and art. The dataset is characterized by requiring models to perform long-distance multi-frame reasoning, which demands not only perceiving the question frame but also analyzing evidence frames distant from it. The tasks include two types: explicit reasoning and implicit reasoning, aiming to evaluate models' capabilities in video understanding, emotion recognition, causal reasoning, sequential structure reasoning, counter-intuitive reasoning, cross-modal transfer reasoning, and video type and intention reasoning, among others. The dataset is constructed following three principles: multi-frame, deep reasoning, and realism. With widely sourced videos and rigorously designed tasks, it aims to promote research on multimodal reasoning capabilities.
MMR-V数据集概述
数据集名称
MMR-V
数据集简介
Can Your MLLMs "Think with Video"? A Benchmark for Multimodal Deep Reasoning in Videos
相关链接




