MMVU
收藏资源简介:
MMVU(Measuring Expert-Level Multidiscipline Video Understanding)是由耶鲁大学NLP团队创建的一个综合性基准数据集,旨在评估多模态基础模型在专家级视频理解任务中的表现。该数据集包含3000个专家标注的问题,覆盖科学、医疗、人文与社会科学、工程四个核心学科的27个主题。每个问题都基于1529个专业领域的视频,要求模型结合领域知识和专家级推理能力进行分析。数据集的创建过程采用了教科书引导的标注方法,确保每个问题都经过严格的质量控制,并附有专家标注的推理过程和相关领域知识。MMVU的应用领域主要集中在专家级知识密集型视频理解任务,旨在解决当前多模态模型在复杂视频理解中的局限性问题。
MMVU (Measuring Expert-Level Multidiscipline Video Understanding) is a comprehensive benchmark dataset developed by the Yale University NLP team, designed to evaluate the performance of multimodal foundation models on expert-level video understanding tasks. This dataset contains 3,000 expert-annotated questions covering 27 topics across four core disciplines: science, medicine, humanities and social sciences, and engineering. Each question is grounded in 1,529 professional-domain videos, requiring models to conduct analysis by integrating domain-specific knowledge and expert-level reasoning abilities. The dataset was constructed using a textbook-guided annotation workflow, with strict quality control applied to every question, and each entry is accompanied by expert-annotated reasoning chains and relevant domain knowledge. The primary application scenarios of MMVU focus on expert-level knowledge-intensive video understanding tasks, aiming to address the limitations of current multimodal models in complex video understanding.




