COVER (COunterfactual VidEo Reasoning)
收藏资源简介:
COVER数据集是由西湖大学工程学院和杭州电子科技大学计算机科学与技术学院联合构建的多维度多模态视频推理基准。该数据集通过引入子问题推理机制,将复杂问题分解为必要条件的多个步骤,对MLLMs的逻辑推理能力进行深入评估。数据集涵盖了多种现实世界场景,包括日常活动识别到复杂场景理解,旨在提高模型在动态和反事实推理任务中的鲁棒性。
The COVER Dataset is a multi-dimensional and multimodal video reasoning benchmark jointly constructed by the School of Engineering of Westlake University and the School of Computer Science and Technology of Hangzhou Dianzi University. By introducing a sub-problem reasoning mechanism, this dataset decomposes complex problems into multiple steps based on necessary conditions to conduct in-depth evaluations of the logical reasoning capabilities of MLLMs. The dataset covers a wide range of real-world scenarios, ranging from daily activity recognition to complex scene understanding, and aims to improve the robustness of models in dynamic and counterfactual reasoning tasks.

- 1Reasoning is All You Need for Video Generalization: A Counterfactual Benchmark with Sub-question Evaluation西湖大学工程学院, 杭州电子科技大学计算机科学与技术学院 · 2025年



