Video-MME - 视频分析多模态大模型评估基准数据集
收藏资源简介:
Video-MME是北京大学、香港大学等6所高校联手,发布的首个专为视频分析设计的多模态大模型评估基准。该数据集包含900个视频,总时长达256小时,研究人员通过反复观看视频内容,手动选择和注释共设计了2,700个高质量的多选题。数据集涵盖6大视觉领域,包括知识、电影与电视、体育竞赛、艺术表演、生活记录和多语言,并进一步细分为天文学、科技、纪录片等30个类别,视频长度从11秒到1小时不等。此外,Video-MME还整合字幕和音频轨道,增强了对视频理解的多模态输入分析。更难能可贵的是,Video-MME中所有数据,包括问答、视频、字幕和音频,都是手工收集和整理的,确保了该基准的高质量。该数据集的创建不仅为研究人员提供了一个富有挑战性的测试基准,也为研究外部信息对视频理解性能的影响提供了宝贵的资源。
Video-MME is the first multimodal large model evaluation benchmark specifically designed for video analysis, jointly released by six universities including Peking University and the University of Hong Kong. The dataset comprises 900 videos with a total duration of 256 hours. Researchers manually selected and annotated 2,700 high-quality multiple-choice questions by repeatedly viewing the video content. The dataset spans six major visual domains, including knowledge, film and television, sports competitions, artistic performances, life recordings, and multilingual content, further subdivided into 30 categories such as astronomy, technology, and documentaries, with video lengths ranging from 11 seconds to 1 hour. Additionally, Video-MME integrates subtitles and audio tracks, enhancing the multimodal input analysis for video comprehension. Notably, all data in Video-MME, including Q&A, videos, subtitles, and audio, were manually collected and curated, ensuring the high quality of this benchmark. The creation of this dataset not only provides researchers with a challenging test benchmark but also offers a valuable resource for studying the impact of external information on video comprehension performance.
数据集概述
名称: Video-MME
描述: Video-MME 是首个全面评估多模态大型语言模型(MLLMs)在视频分析中应用的基准数据集。该数据集旨在全面评估 MLLMs 处理视频数据的能力,涵盖广泛的视觉领域、时间持续性和数据模态。
数据集构成:
- 视频数量: 900 个
- 总时长: 254 小时
- 问题-答案对: 2,700 个人工标注的问答对
数据集特点:
- 时间维度持续性: 包括短(<2分钟)、中(4分钟~15分钟)和长(30分钟~60分钟)视频,范围从11秒到1小时。
- 视频类型多样性: 涵盖6个主要视觉领域,包括知识、电影与电视、体育竞赛、生活记录和多语言,共有30个子领域。
- 数据模态广度: 除视频帧外,还包括字幕和音频,以评估 MLLMs 的全方位能力。
- 标注质量: 所有数据均为新收集并由人工标注,确保多样性和质量。
使用许可:
- 仅限学术研究使用,禁止任何形式的商业使用。
- 所有视频的版权属于视频所有者。
- 未经事先批准,不得以任何方式分发、发布、复制、传播或修改 Video-MME 的全部或部分内容。
评估流程:
- 提取帧和字幕: 包括900个视频和744个字幕,所有长视频均包含字幕。
- 评估方法: 使用特定的 JSON 格式记录模型响应,并通过自定义脚本计算准确率。
联系方式: 如有任何问题,请发送邮件至 videomme2024@gmail.com。




