CodecFake-Omni
收藏资源简介:
CodecFake-Omni是由国立台湾大学等机构创建的大规模数据集,旨在研究基于神经编解码器的深度伪造语音检测。该数据集包含31种不同的开源编解码器模型生成的训练数据,以及17种先进的CoSG模型生成的测试数据。数据集通过重新合成真实语音生成训练数据,测试数据则来自未发布的模型生成的语音。CodecFake-Omni是目前最大的CodecFake语料库,涵盖了最广泛的编解码器架构。该数据集的应用领域主要是深度伪造语音检测,旨在解决当前反欺骗模型在检测由CoSG系统生成的合成语音时的不足问题。
CodecFake-Omni is a large-scale dataset developed by institutions including National Taiwan University for research on neural codec-based deepfake speech detection. This dataset comprises training data generated by 31 distinct open-source codec models, as well as test data produced by 17 state-of-the-art CoSG models. The training data is generated via resynthesis of real speech, while the test data is sourced from speech produced by unreleased models. CodecFake-Omni currently stands as the largest CodecFake corpus, covering the broadest range of codec architectures. The primary application scenario of this dataset is deepfake speech detection, with the goal of addressing the limitations of current anti-spoofing models when detecting synthetic speech generated by CoSG systems.

- 1CodecFake-Omni: A Large-Scale Codec-based Deepfake Speech Dataset国立台湾大学 · 2025年



