ParsVoice
收藏资源简介:
ParsVoice是一个大规模的多说话人波斯语音语料库,专为文本到语音(TTS)合成应用而设计。该数据集包含3526小时的语音,经过筛选后形成了1804小时的高质量子集,拥有超过470个说话人。ParsVoice是迄今为止最大的高质量波斯语音数据集,提供了与主要英语语料库相当的说话人多样性和音频质量。数据集的创建过程包括数据收集、智能音频分割、边界优化算法、多维度质量评估以及说话人识别等步骤。ParsVoice旨在促进波斯语音技术的发展,并为其他低资源语言提供一个模板。
ParsVoice is a large-scale multi-speaker Persian speech corpus designed specifically for text-to-speech (TTS) synthesis applications. This corpus contains 3526 hours of raw speech audio, and a filtered high-quality subset of 1804 hours is derived, comprising over 470 unique speakers. ParsVoice is the largest high-quality Persian speech dataset to date, offering speaker diversity and audio quality comparable to major English-language speech corpora. The construction process of ParsVoice includes multiple procedures such as data collection, intelligent audio segmentation, boundary optimization algorithms, multi-dimensional quality evaluation, and speaker identification. ParsVoice aims to promote the development of Persian speech technologies and serve as a template for other low-resource languages.

- 1通过伊朗德黑兰大学电气与计算机工程学院 · 2025年



