NileTTS
收藏资源简介:
NileTTS是由乔治亚理工学院和尼罗大学联合构建的首个公开埃及阿拉伯语语音合成数据集,包含38小时双说话人(男女各一)的语音数据,覆盖医疗、销售及日常对话三大领域。该数据集通过大语言模型生成埃及阿拉伯语文本,经神经音频合成工具转换为自然语音后,采用Whisper自动转录并经过人工质检,最终形成9521条高质量语音-文本对。其创新性的合成流程为低资源方言语音研究提供了可扩展方案,主要应用于改进埃及阿拉伯语的TTS模型训练,解决该方言在语音助手等场景下的技术空白问题。
NileTTS is the first publicly available Egyptian Arabic speech synthesis dataset jointly developed by the Georgia Institute of Technology and Nile University. It contains 38 hours of speech data from two speakers (one male and one female), covering three domains: medical, sales and daily conversations. The dataset is created by first generating Egyptian Arabic texts via large language models (LLMs), converting them into natural speech using neural audio synthesis tools, then automatically transcribing the speech with Whisper and conducting manual quality checks, ultimately resulting in 9521 high-quality speech-text pairs. Its innovative synthesis pipeline provides a scalable solution for low-resource dialect speech research, and it is mainly applied to improve the training of Egyptian Arabic TTS models, addressing the technical gap in scenarios such as voice assistants for this dialect.



