MINERVA
收藏资源简介:
MINERVA是一个用于现代多模态模型的新型视频推理数据集。每个问题都附带5个答案选项以及详细的、手工制作的推理轨迹。数据集是多模态的,视频领域和长度多样化,包含复杂的多步问题。广泛的基准测试表明,我们的数据集对前沿开源和专有模型提出了挑战。我们进行了细粒度的错误分析,以确定各种模型中的常见失败模式,并创建了一个推理错误的分类法。我们使用这个分类法来探索人类和LLM-asa-judge方法对视频推理轨迹的评分,并发现失败模式主要与时间定位相关,其次是视觉感知错误。数据集、问题、答案候选和推理轨迹将在https://github.com/googledeepmind/neptune?tab=readme-ov-file#minerva公开提供。
MINERVA is a novel video reasoning dataset tailored for modern multimodal models. Each question is accompanied by 5 answer options and detailed, handcrafted reasoning traces. The dataset is multimodal, with diverse video domains and lengths, and includes complex multi-step questions. Extensive benchmark evaluations demonstrate that our dataset poses significant challenges to state-of-the-art open-source and proprietary models. We conducted fine-grained error analysis to identify common failure patterns across various models, and developed a taxonomy of reasoning errors. Using this taxonomy, we explored both human and LLM-as-a-judge methods for scoring video reasoning traces, and found that failure modes are primarily associated with temporal localization, followed by visual perception errors. The dataset, questions, answer candidates, and reasoning traces will be publicly available at https://github.com/googledeepmind/neptune?tab=readme-ov-file#minerva.




