遇见数据集

STOPA: A Dataset of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution

收藏
Zenodo2025-06-06 更新2026-05-26 收录
官方服务:

资源简介:

STOPA is a dataset for source tracing and attribution of deepfake audio, identifying which synthesis system generated a given utterance. It includes over 700,000 synthetic speech samples, generated using 13 distinct systems with controlled variation across 8 acoustic models and 6 vocoders. The dataset follows ASVspoof2019 Logical Access protocols and uses speakers from the VCTK corpus. It supports open-world evaluation, where test utterances are compared against individual source hypotheses without assuming closed-set conditions. Rich metadata and pairwise trial protocols enable fine-grained attribution at the level of attack, acoustic model, or vocoder. All audio is provided as 16-bit PCM WAV at 16 kHz. Metadata includes transcription, silence regions, WER, and system labels. License: CC BY 4.0Audio type: Synthetic onlyLanguage: English (VCTK-based)

提供机构:
Zenodo
创建时间:
2025-05-26
二维码
社区交流群
二维码
科研交流群
商业服务