ML-DF: A Proof-of-Concept EU AI Act Compliance Supporting Deepfake Speech Dataset
收藏资源简介:
ML-DF - A Multilingual Deepfake Speech Dataset for Fairness and Compliance Analyses ML-DF is a proof-of-concept dataset for evaluating deepfake (synthetic) speech detectors with a focus on fairness, auditability, and EU AI Act–aligned data governance. It contains 448,000 utterances (balanced 224,000 bona fide + 224,000 deepfake) designed to support subgroup analysis and compliance-oriented audits (e.g., by gender and language). Most samples are 2–6 seconds long. The dataset includes structured metadata (language, speaker info, synthesis tool) to enable fine-grained analyses. The dataset is balanced and targets realistic attack coverage by combining modern TTS and VC systems. Contents Total size: 448,000 utterances (224k genuine, 224k deepfake) Languages: English, German, French, Spanish, Italian Speakers: 106 total (53 male + 53 female) Synthesizers TTS: VITS [1], ZMM-TTS [2] VC: LVC-VC [3], DDDM-VC [4] Typical duration: 2–6 s per utterance Balance: Balanced across real/fake, gender, and languages; convenient for statistical evaluation All audio files are provided in 16-bit PCM WAV format at 16 kHz, consisting of synthetic (deepfake) and genuine speech samples. Source & Metadata Source speech is derived from Multilingual Librispeech (MLS) [5]. Each speaker contributed 3 hours of speech for model training, ensuring sufficient data for high-quality synthesis. For the development set, 70 minutes of speech per speaker were used. For the test set, another 70 minutes of speech per speaker were utilized. All subsets were speaker-disjoint to prevent data leakage and allow fair model evaluation - the data used for training the synthesizers is different from the ones used for synthesis, following the train-dev-test split of the MLS corpus. Therefore, ML-DF provides speaker-disjoint genuine vs. spoofed content and structured metadata (language, synthesis method details, reference speaker information). The language subset archives contain protocols with per-recording metadata and demographics. Details about the synthesis procedure (including the training pipeline) are available separately in the file synthesis.pdf. All code for segmenting and aligning data, training synthesizers, generating deepfake recordings, as well as training and evaluating detectors, is available in the code.zip archived repository. To prevent shortcut learning and inflating detector performance [6], we measured several auxiliary properties (artefacts) of recordings in ML-DF. We used the measured values as simple detector scores to compute Equal Error Rates (EER) to assess their discriminative power. We consider an artefact sufficiently controlled when its distribution between bona fide and deepfake samples largely overlaps, i.e., EER ≥ 30%. The results of this analysis are in the file artefacts.pdf. License CC BY 4.0 (commercial use permitted). The dataset builds upon MLS (also CC BY 4.0) and is released under the same license. References [1] J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” in Proceedings of the 38th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul 2021, pp. 5530–5540. [Online]. Available: https://proceedings.mlr.press/v139/kim21f.html [2] C. Gong, X. Wang, E. Cooper, D. Wells, L. Wang, J. Dang, K. Richmond, and J. Yamagishi, “Zmm-tts: Zero-shot multilingual and multispeaker speech synthesis conditioned on self-supervised discrete speech representations,” vol. 32, p. 4036–4051, Sep. 2024. [Online]. Available: https://doi.org/10.1109/TASLP.2024.3451951 [3] W. Kang, M. Hasegawa-Johnson, and D. Roy, “End-to-end zero-shot voice conversion with location-variable convolutions,” in Interspeech 2023, 2023, pp. 2303–2307 [4] H.-Y. Choi, S.-H. Lee, and S.-W. Lee, “Dddm-vc: Decoupled denoising diffusion models with disentangled representation and prior mixup for verified robust voice conversion,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 16, 2024, pp. 17 862–17 870. [5] V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “MLS: A large-scale multilingual dataset for speech research,” in Interspeech 2020, 2020, pp. 2757–2761 [6] N. Müller, F. Dieckmann, P. Czempin, R. Canals, K. Böttinger, and J. Williams, “Speech is silver, silence is golden: What do asvspoof-trained models really learn?” in 2021 Edition of the Automatic Speaker Verification and Spoofing Countermeasures Challenge, 2021, pp. 55–60.



