SwanBench-Speech
收藏资源简介:
SwanBench-Speech是由浙江大学与字节跳动联合构建的综合性长语音生成基准数据集,旨在系统评估模型在多样化场景下的表现。该数据集包含1,101个测试样本,覆盖17种下游语音场景,涵盖声学、语义和表达力三大核心挑战,数据来源于在线文本语料库、音频媒体及大语言模型生成,并经过严格的去重、质量过滤和人工校验流程。该数据集主要应用于长文本语音合成和对话生成领域,旨在解决现有评估方法在场景覆盖度、一致性及表达力维度上的不足,为模型性能提供标准化、细粒度的自动化评估框架。
SwanBench-Speech is a comprehensive long-form speech generation benchmark dataset jointly developed by Zhejiang University and ByteDance, aiming to systematically evaluate model performance across diverse scenarios. This dataset contains 1,101 test samples, covering 17 downstream speech scenarios and encompassing three core challenges: acoustics, semantics, and expressiveness. The data is sourced from online text corpora, audio media, and outputs generated by large language models (LLMs), and has undergone strict deduplication, quality filtering, and manual verification processes. It is primarily applied in the fields of long-text speech synthesis and dialogue generation, aiming to address the shortcomings of existing evaluation methods in terms of scenario coverage, consistency and expressiveness, and provide a standardized, fine-grained automated evaluation framework for model performance.
数据集名称
SwanBench-Speech
所属机构
ByteDance(字节跳动)
发布会议/期刊
ACL 2026
基准概述
SwanBench-Speech 是一个用于评估长文本语音生成(Long-Form Speech Generation)的综合基准,覆盖多样化场景。其核心目标在于系统性地评估模型在长上下文条件下的表现,弥补现有评估在场景覆盖度和长文本因素(如一致性与连贯性)上的不足。
关键特性
- 丰富的语音场景:聚焦长文本语音生成与对话生成,涵盖声学、语义和表现力挑战,包含 1,101 个样本,覆盖 17 个常见语音场景。
- 全面的评估维度:沿声学、语义和表现力三个轴,定义了一个包含七项指标的自动化评估协议,以提供全面、准确、标准化的评估。
- 有价值的洞察:通过大量实验揭示,当前模型在高表现力场景中仍存在困难,且与真实录音相比,在一致性和层次性方面存在显著差距。
演示示例
- 按原始基准部分划分为:按维度(Per-Dimensions)、按场景(Per-Scenarios)和消融研究(Ablation Study)。
- 示例涵盖语音(Speech)和对话(Dialog)样本,并展示了多个模型的评估结果,如 GLM-TTS、ZipVoice、FishSpeech、Minimax、SparkTTS、Seed-TTS、VibeVoice、Gemni、IndexTTS2、F5TTS、CosyVoice3、MegaTTS3、CosyVoice2、Elenvenlabs、InworldTTS、OpenAI、MOSS-TTSD、FireRedTTS2、MoonCast、Seed-TTS-Podcast、SoulX-Podcast、ZipVoice Dialogue 等。





