VoxENES 2026
收藏资源简介:
VoxENES 2026是由南佛罗里达大学创建的现代双语语音欺骗检测基准数据集,旨在评估深度学习模型在LLM驱动的语音合成技术下的泛化能力。该数据集包含53,628个音频样本,涵盖英语和西班牙语,数据来源包括LibriSpeech和VoxPopuli的真实语音,以及10种当代TTS和VC系统生成的合成音频,并应用了10种标准化后处理增强以模拟真实传输条件。数据集的构建过程涉及真实语音标准化、多样化合成系统选择以及系统性后处理扰动。该数据集主要应用于音频深度伪造检测领域,旨在解决现有检测器因数据漂移问题而导致的泛化性能不足,为开发鲁棒的语音欺骗对抗措施提供实用测试平台。
VoxENES 2026 is a modern bilingual speech spoofing detection benchmark dataset developed by the University of South Florida, designed to evaluate the generalization capability of deep learning models against LLM-driven speech synthesis technologies. This dataset contains 53,628 audio samples encompassing English and Spanish, with data sources including authentic speech from LibriSpeech and VoxPopuli, as well as synthesized audio generated by 10 contemporary Text-to-Speech (TTS) and Voice Conversion (VC) systems. Ten standardized post-processing augmentations are applied to simulate real-world transmission conditions. The construction of this dataset involves authentic speech normalization, selection of diverse synthesis systems, and systematic post-processing perturbations. This dataset is primarily applied in the field of audio deepfake detection, aiming to address the insufficient generalization performance of existing detectors caused by data drift issues, and provide a practical testbed for developing robust anti-spoofing countermeasures against speech spoofing.
数据集概述:VoxENES 2026
- 名称:VoxENES 2026
- 用途:用于评估针对现代LLM时代的文本转语音(TTS)和语音转换(VC)系统的音频深度伪造检测器。
- 语言:双语(英语/西班牙语)
- 规模:共 53,628 个音频样本
- 真实语音(Bonafide):3,028 个样本
- 原始合成样本:4,600 个样本(来自10种合成方法:7种TTS + 3种VC)
- 增强变体:46,000 个样本(经过10种后处理条件下的增强)
数据结构
bonafide/english/— 1,500个真实英语语音样本(来自LibriSpeech)bonafide/spanish/— 1,528个真实西班牙语语音样本(来自VoxPopuli)tts/original/{method}/{language}/{fixed,random}/tts/augmented/{method}/{language}/{fixed,random}/{augmentation}/vc/original/{method}/{language}/vc/augmented/{method}/{language}/{augmentation}/
合成系统
- TTS(文本转语音):VoxCPM 1.5, Qwen3-TTS, GLM-TTS, FlashLabs Chroma, VibeVoice, CosyVoice 3, Chatterbox ML
- VC(语音转换):Seed-VC, OpenVoice v2, RVC v2
增强处理
MP3 64k, AAC 128k, 白噪声(10/20 dB), 婴儿噪声(15 dB), 重采样(8k/16k), 速度扰动(0.9x/1.1x), 音量归一化
许可协议
- 许可证类型:Attribution 4.0 International (CC BY 4.0)
- 预期更新频率:未指定
其他信息
- 标签:Audio, Linguistics, Audio Classification, Text-To-Speech, Speech Synthesis
- 数据集 ID:
interspeech_2712/voxenes-2026 - 版本:Version 1(大小:23.33 GB,文件数:53.6k)

- 1VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion南佛罗里达大学 · 2026年



