OmniEval
收藏资源简介:
OmniEval是一个全面的基准数据集,旨在评估能够同时处理视觉、听觉和文本信息的全模式模型。数据集包含810个音视频同步视频片段,包括285个中文视频和525个英文视频,以及2617个问答对,涵盖1412个开放式问题和1205个多项选择题。OmniEval设计了一系列强调音频和视频强耦合的任务,要求模型有效地利用所有模态的协作感知。数据集通过结合自动化数据处理和人工审核的方式创建,旨在提供一个具有挑战性和可靠性的资源,用于评估Omni模型在多种认知任务中的能力,包括细粒度事件理解。OmniEval旨在解决现有全模式模型评估的不足,推动全模式模型的发展,促进研究人员构建更强大的模型,能够理解和构建所有模态的上下文中的连贯性。
OmniEval is a comprehensive benchmark dataset designed to evaluate full-modal models that can simultaneously process visual, auditory and textual information. The dataset comprises 810 synchronized audio-visual video clips, including 285 Chinese-language videos and 525 English-language videos, alongside 2617 question-answer pairs encompassing 1412 open-ended questions and 1205 multiple-choice questions. OmniEval features a suite of tasks that highlight the strong coupling between audio and visual modalities, requiring models to effectively leverage collaborative perception across all modalities. Developed through a combination of automated data processing and manual review, this dataset aims to offer a challenging and reliable resource for evaluating the capabilities of full-modal models across a wide range of cognitive tasks, including fine-grained event understanding. OmniEval seeks to address the limitations of existing full-modal model evaluations, advance the development of full-modal models, and empower researchers to construct more robust models that can comprehend and build coherence within cross-modal contextual settings.
OmniEval 数据集概述
基本信息
- 数据集名称: OmniEval
- 开发团队:
- Yiman Zhang, Ziheng Luo, Qiangyu Yan, Wei He, Borui Jiang, Xinghao Chen, Kai Han
- 机构: Huawei Noah’s Ark Lab, University of Science and Technology of China
- 联系方式:
- yiman.zhang@huawei.com
- xinghao.chen@huawei.com
- kai.han@huawei.com
数据集特点
- 多模态输入: 视觉、听觉和文本输入
- 全模态协作: 设计强调音频和视频之间强耦合的评估任务
- 视频多样性:
- 810个音频-视觉同步视频
- 285个中文视频
- 525个英文视频
- 任务多样性:
- 2617个问答对
- 1412个开放式问题
- 1205个多项选择题
- 3种主要任务类型
- 12种子任务类型
- 引入细粒度视频定位任务(Grounding)
- 2617个问答对
实验与结果
- 评估模型: 多个全模态模型
- 主要发现: 现有模型在理解真实世界信息方面面临显著挑战
相关资源
- 论文: arXiv链接
- 项目页面与数据: OmniEval项目页面

- 1OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs华为诺亚方舟实验室 · 2025年



