FMFCC-A
收藏资源简介:
FMFCC-A数据集是由中国科学院信息工程研究所信息安全国家重点实验室创建的一个挑战性普通话数据集,专注于合成语音检测。该数据集包含50,000条普通话语音,其中40,000条为合成语音,由11个普通话TTS系统和两个普通话VC系统生成,10,000条为真实语音,来自58位不同年龄和性别的说话者。数据集分为训练、开发和评估集,用于研究在未知语音合成系统或音频后处理操作下的合成普通话语音检测。该数据集旨在填补普通话合成语音检测数据集的空白,并推动相关技术的发展。
FMFCC-A dataset is a challenging Mandarin speech dataset curated by the State Key Laboratory of Information Security, Institute of Information Engineering, Chinese Academy of Sciences, focusing on synthetic speech detection. This dataset contains 50,000 Mandarin speech samples, among which 40,000 are synthetic speech generated by 11 Mandarin TTS systems and 2 Mandarin VC systems, while the remaining 10,000 are genuine speech samples from 58 speakers with diverse ages and genders. The dataset is split into training, development and evaluation subsets to support research on synthetic Mandarin speech detection under unknown speech synthesis systems or audio post-processing operations. This dataset is designed to fill the gap in existing Mandarin synthetic speech detection datasets and promote the development of related technologies.

- 1FMFCC-A: A Challenging Mandarin Dataset for Synthetic Speech Detection中国科学院信息工程研究所信息安全国家重点实验室 · 2021年



