LVOmniBench
收藏资源简介:
LVOmniBench是由浙江大学、西湖大学等机构联合创建的首个专注于长音频-视频跨模态理解的基准数据集。该数据集包含275个高质量长视频,时长介于10至90分钟,总时长140小时,覆盖娱乐、生活方式等五大领域,并严格筛选动态视听内容。通过人工标注构建了1,014个需跨模态推理的多选题,涵盖感知、理解、推理和逻辑四个认知维度。该数据集旨在解决现有评估对长视频(如影视、纪录片)理解不足的问题,推动全模态模型在复杂时空对齐和细粒度推理方面的研究。
LVOmniBench is the first benchmark dataset dedicated to long audio-visual cross-modal understanding, jointly created by institutions including Zhejiang University, Westlake University, and other relevant organizations. This dataset includes 275 high-quality long videos, with individual durations ranging from 10 to 90 minutes and a total cumulative duration of 140 hours, covering five major domains such as entertainment and lifestyle, with strictly curated dynamic audio-visual content. It features 1,014 multiple-choice questions requiring cross-modal reasoning, constructed via manual annotation, which cover four cognitive dimensions: perception, understanding, reasoning, and logic. This dataset is designed to address the gap in existing evaluations that lack sufficient assessment of long video (e.g., films and documentaries) understanding, and advance research on full-modal models in terms of complex spatio-temporal alignment and fine-grained reasoning.
LVOmniBench 数据集概述
数据集基本信息
- 数据集名称: LVOmniBench
- 核心目标: 为全能模态大语言模型(OmniLLMs)在长音频-视频理解评估方面提供开创性的综合评估基准。
- 解决的问题: 当前评估主要针对10秒至5分钟的短音频和视频片段,无法反映现实应用中通常长达数十分钟的视频理解需求。
数据集构成
- 视频数量: 275个高质量视频。
- 视频时长范围: 10至90分钟。
- 视频平均时长: 2,069秒。
- 问答对数量: 1,014个手动构建的问答对。
- 问答对设计: 明确要求对音频和视觉模态进行联合推理。
评估维度与结果
- 难度等级: 低(Low)、中(Medium)、高(High)。
- 能力维度: 理解(Understanding)、感知(Perception)、推理(Inference)、逻辑(Logical)。
- 评估模型示例:
- Gemini-3.0-Pro (A + V): 平均得分65.8
- Gemini-3.0-Flash (A + V): 平均得分59.0
- Qwen3-VL-30B (V): 平均得分36.3
- Qwen2-Audio (A): 平均得分24.7
引用信息
-
引用格式:
@article{tao2026lvomnibench, title={LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs}, author={Keda Tao and Yuhua Zheng and Jia Xu and Wenjie Du and Kele Shao and Hesong Wang and Xueyi Chen and Xin Jin and Junhan Zhu and Bohan Yu and Weiqiang Wang and Jian Liu and Can Qin and Yulun Zhang and Ming-Hsuan Yang and Huan Wang}, journal={arXiv preprint arXiv:2603.19217}, year={2026} }

- 1LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs浙江大学; 西湖大学; 蚂蚁集团; 上海创新研究院; 上海交通大学; 加州大学默塞德分校 · 2026年



