遇见数据集

System Fingerprint Recognition for Deepfake Audio (SFR) - Clean Set

收藏
Zenodo2024-08-20 更新2026-05-26 收录
官方服务:

资源简介:

The rapid progress of deep speech synthesis models has posed significant threats to society such as malicious manip ulation of content. This has led to an increase in studies aimed at detecting so-called “deepfake audio”. However, existing works focus on the binary detection of real audio and fake audio. In real-world scenarios such as model copyright protection and digital evidence forensics, it is needed to know what tool or model generated the deepfake audio to explain the decision. This motivates us to ask: ‘Can we recognize the system fingerprints of deepfake audio?’ In this paper, we present the first deepfake audio dataset for System Fingerprint Recognition (SFR) and conduct an initial investigation. We collected the dataset from the speech synthesis systems of seven Chinese vendors that use the latest state-of-the-art deep learning technologies, including both clean and compressed sets. In addition, we provide extensive benchmarks and research findings to facilitate the further development of system fingerprint recognition methods. The dataset is publicly available. The subsets 01, 02, and 03 represent the training set, development set, and test set, respectively. This data set is licensed with a CC BY-NC-ND 4.0 license.

深度语音合成模型的飞速发展,已对社会造成诸多严峻威胁,例如针对内容的恶意篡改。这使得针对所谓“深度伪造音频”的检测研究数量显著增长。然而,现有研究多聚焦于真实音频与伪造音频的二元分类检测。在模型版权保护、数字证据取证等真实应用场景中,往往需要明确深度伪造音频的生成工具或模型,以对检测结果作出合理解释。为此我们提出研究问题:能否识别深度伪造音频的系统指纹? 本文构建了首个面向系统指纹识别(System Fingerprint Recognition,SFR)的深度伪造音频数据集,并开展了初步探索性研究。本数据集采集自7家采用当前最先进深度学习技术的中国厂商的语音合成系统,包含纯净音频与压缩音频两类子集。此外,我们还提供了全面的基准测试与研究结论,以推动系统指纹识别相关方法的后续发展。本数据集已公开发布。 子集01、02、03分别对应训练集、开发集与测试集。 本数据集采用CC BY-NC-ND 4.0协议进行授权。

提供机构:
Zenodo
创建时间:
2024-08-19
二维码
社区交流群
二维码
科研交流群
商业服务