WavBench
收藏资源简介:
WavBench是由浙江大学等机构联合开发的端到端语音对话模型综合评测基准,旨在解决现有评测标准在认知复杂性、口语化表达及副语言特征方面的不足。该数据集包含三个核心子集:Pro子集针对高难度推理任务设计,Basic子集定义口语自然度的新标准,Acoustic子集覆盖10种副语言特征(如年龄、情感、背景音等)。数据通过GPT-4等大语言模型与商用级TTS工具生成,涵盖数学、编程、安全等7大认知领域,适用于评估复杂推理与情感交互场景下的语音模型表现。
WavBench is a comprehensive evaluation benchmark for end-to-end speech dialogue models, jointly developed by Zhejiang University and other institutions. It aims to address the limitations of existing evaluation standards in terms of cognitive complexity, colloquial expressions and paralinguistic features. This dataset includes three core subsets: the Pro subset is tailored for high-difficulty reasoning tasks, the Basic subset establishes a new benchmark for spoken naturalness, and the Acoustic subset covers 10 paralinguistic features such as age, emotion, background audio and so on. The data is generated using large language models (e.g., GPT-4) and commercial-grade text-to-speech (TTS) tools, covering seven cognitive domains including mathematics, programming, cybersecurity and other fields. It is suitable for evaluating the performance of speech models in scenarios involving complex reasoning and emotional interaction.
WavBench 数据集概述
数据集名称
WavBench
核心目标
评估端到端语音对话模型在真实世界场景中的能力,重点关注音频中心的口语语义和副语言保真度。
基准框架
- Pro 子集:用于挑战具备推理能力的模型,包含复杂和判别性任务。
- Basic 子集:用于基准测试口语适应性,定义了一个以“可听性”为核心的新标准。
- Acoustic 子集:用于评估全面的副语言交互能力。
数据集规模
- 包含 17,577 个项目。
- 总时长 76.5 小时。
评估维度与内容
1. 口语表达
- 认知领域:涵盖创意写作、指令遵循、代码、数学、问答、安全性和逻辑共7个领域。
- 组织层级:分为 Basic 和 Pro 两个层级。
2. 声学交互
- 组成部分:包含显式理解、显式生成和隐式对话。
- 评估的副语言维度(10个):
- 说话者信息:年龄、性别、口音、语言。
- 声学特征:音高、语速、音量、情感。
- 背景声音:音频事件、音乐。
- 具体属性:
- 年龄:儿童、青少年、中年、老年。
- 性别:男性、女性。
- 口音:印度、加拿大、英国、新加坡、美国、澳大利亚。
- 语言:中文、英文。
- 音高:低、正常、高。
- 语速:慢、正常、快。
- 音量:低、正常、高。
- 情感:中性、快乐、悲伤、愤怒、惊讶、厌恶、恐惧。
- 音频事件:风声、人群声、雷声、发令枪声、摔门声。
- 音乐:钢琴、吉他、鼓。
3. 交互类型
- 评估显式指令和隐式多轮对话,要求模型在没有直接提示的情况下推断声学线索。
评估模型
评估了五款先进的端到端语音对话模型:
- Qwen3-Omni
- Kimi-Audio
- Mimo-Audio
- Step-Audio-2
- GPT-4o Audio
评估结果概览
评估结果分为五个面板:
- 面板A:口语表达能力 - Pro子集
- 面板B:口语表达能力 - Basic子集
- 面板C:声学显式理解
- 面板D:声学显式生成
- 面板E:隐式声学交互能力
引用信息
- 标题:WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models
- 作者:Yangzhuo Li, Shengpeng Ji, Yifu Chen, Tianle Liang, Haorong Ying, Yule Wang, Junbo Li, Jun Fang, Zhou Zhao
- 年份:2026
- arXiv ID:2602.12135
- arXiv 链接:https://arxiv.org/abs/2602.12135

- 1WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models厦门大学; 浙江大学; 香港中文大学·深圳 · 2026年



