ArVoice
收藏资源简介:
ArVoice是一个包含多种发音的现代标准阿拉伯语(MSA)语音语料库,包含带变音符号的转录文本,主要用于多发音人语音合成,也可用于语音基转换、声音转换和深度伪造检测等任务。该数据集包括:1)六位声音人才的新专业录音集,具有多样化的人口统计数据;2)阿拉伯语音语料库的修改子集;3)来自两个商业系统的优质合成语音。整个语料库共有83.52小时的语音,涵盖11个声音,其中约10小时由7位说话者的真人声音组成。我们训练了三个开源的文本到语音(TTS)和两个声音转换系统,以展示数据集的使用案例。语料库可供研究使用。
ArVoice is a speech corpus for Modern Standard Arabic (MSA) featuring diverse pronunciations, equipped with diacritized transcriptions. Primarily designed for multi-speaker text-to-speech synthesis, it can also be utilized for tasks such as speech-based conversion, voice conversion, and deepfake detection. The dataset comprises three components: 1) A new professional recording dataset from six voice talents with diverse demographic characteristics; 2) A modified subset of an existing Arabic speech corpus; 3) High-quality synthesized speech from two commercial systems. The entire corpus spans 83.52 hours of speech covering 11 voice profiles, with approximately 10 hours originating from the real human voices of 7 speakers. We have trained three open-source text-to-speech (TTS) systems and two voice conversion systems to demonstrate the practical use cases of this corpus. The corpus is available for research purposes.

- 1ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Mohamed Bin Zayed University of Artificial Intelligence, UAE · 2025年



