遇见数据集

PhonemeDF: A Synthetic Speech Dataset for Audio Deepfake Detection and Naturalness Evaluation

收藏
Zenodo2026-04-12 更新2026-05-26 收录
官方服务:

资源简介:

PhonemeDF is a large-scale phoneme-level parallel dataset of real and synthetic speech (approximately 730 hours), designed for audio deepfake detection and speech naturalness evaluation. The dataset consists of real speech samples derived from a subset of the LibriSpeech corpus (train-clean-100) and corresponding synthetic speech generated using four Text-to-Speech (TTS) systems (MeloTTS, XTTS v2, Chatterbox TTS, and VITS) and three Voice Conversion (VC) systems (Chatterbox VC, FreeVC, and StarGAN VC). Each audio sample is paired with phoneme-level alignments obtained using the Montreal Forced Aligner (MFA) with the ARPAbet phoneme set. The dataset contains 28,539 real utterances and 199,773 synthetic utterances, totaling approximately 730 hours of speech, along with corresponding TextGrid files. All audio is standardized to 16 kHz. RESOURCES Code and additional resources:https://github.com/Vamshi-Nallaguntla/PhonemeDF Paper (arXiv):https://doi.org/10.48550/arXiv.2603.15037

提供机构:
Zenodo
创建时间:
2026-04-12
二维码
社区交流群
二维码
科研交流群
商业服务