MLVU
收藏资源简介:
MLVU(多任务长视频理解基准)是由北京人工智能研究院创建的一个全面评估长视频理解性能的数据集。该数据集包含2593个任务,涉及多种视频类型,如电影、监控录像、第一人称视频等,覆盖从3分钟到2小时不等的视频长度。数据集的创建过程涉及从多个来源收集视频,并手动标注相关任务。MLVU旨在通过多样化的评估任务,全面测试多模态大型语言模型在长视频理解方面的能力,解决现有基准在视频长度、类型和任务多样性方面的不足。
MLVU (Multi-Task Long Video Understanding Benchmark) is a comprehensive dataset developed by the Beijing Academy of Artificial Intelligence for evaluating long-form video understanding performance. This dataset contains 2,593 tasks, covering diverse video types such as feature films, surveillance footage, first-person videos, and more, with video durations ranging from 3 minutes to 2 hours. The construction of MLVU involved collecting videos from multiple sources and manually annotating the corresponding tasks. The core goal of MLVU is to comprehensively test the long-form video understanding capabilities of multimodal large language models through diverse evaluation tasks, thereby addressing the shortcomings of existing benchmarks in terms of video length range, video category diversity, and task variety.

- 1MLVU: A Comprehensive Benchmark for Multi-Task Long Video Understanding北京人工智能研究院 · 2024年



