Diffusion-Based Synthetic Speech Dataset (DiffSSD)
收藏资源简介:
Diffusion-Based Synthetic Speech Dataset (DiffSSD) 是由普渡大学和米兰理工大学联合创建的一个用于语音取证的扩散模型合成语音数据集。该数据集包含约200小时的标记语音,包括由8个开源和2个商业扩散模型生成的合成语音。数据集内容包括70,000个合成语音信号,涵盖11个不同的说话者,平均语音时长约为7.49秒。数据集的创建过程包括文本生成、语音合成和数据集分割,旨在解决现有合成语音检测方法在检测最新扩散模型生成语音时的不足。DiffSSD的应用领域主要集中在合成语音检测,特别是在防止合成语音的恶意使用方面。
Diffusion-Based Synthetic Speech Dataset (DiffSSD) is a synthetic speech dataset designed for speech forensics, jointly developed by Purdue University and Politecnico di Milano. It contains approximately 200 hours of labeled speech, generated by 8 open-source and 2 commercial diffusion models. The dataset includes 70,000 synthetic speech signals, covering 11 distinct speakers, with an average duration of about 7.49 seconds per signal. The dataset construction workflow encompasses text generation, speech synthesis, and dataset partitioning, aiming to address the shortcomings of existing synthetic speech detection methods when detecting speech produced by state-of-the-art diffusion models. The primary application scenarios of DiffSSD focus on synthetic speech detection, particularly in preventing the malicious misuse of synthetic speech.




