mohammedaly22/lahgtna-levantine-tts
收藏资源简介:
Lahgtna Levantine TTS 是一个合成黎凡特阿拉伯语和英语代码切换的语音数据集,专为文本到语音(TTS)任务设计。数据集包含50,000个话语,由10位说话者(5男5女)生成,总时长约66.8小时。其中,44,154个话语为纯黎凡特阿拉伯语,5,846个话语为黎凡特阿拉伯语与英语的代码切换。数据基于GU-CLASP Shami语料库的真实黎凡特阿拉伯语文本(涵盖叙利亚、黎巴嫩、巴勒斯坦和约旦方言)以及合成的代码切换模板生成。文本经过规范化处理,包括Unicode清理、数字口头化、部分加注音(仅针对同形异义词)和黎凡特词典覆盖。语音合成使用Lahgtna-OmniVoice模型,这是一个针对黎凡特阿拉伯语方言微调的零样本TTS模型,通过参考音频进行语音克隆生成。数据集以24,000 Hz的采样率存储音频,并包含文本转录、说话者ID、说话者姓名、性别和句子类型等字段。
Lahgtna Levantine TTS is a synthetic Levantine Arabic and English code-switching speech dataset designed for text-to-speech (TTS) tasks. The dataset consists of 50,000 utterances generated by 10 speakers (5 male, 5 female), with an estimated total duration of approximately 66.8 hours. It includes 44,154 pure Levantine Arabic utterances and 5,846 code-switching utterances mixing Levantine Arabic and English. The data is derived from the GU-CLASP Shami corpus, which provides authentic Levantine Arabic text from Syrian, Lebanese, Palestinian, and Jordanian sub-dialects, along with synthetic code-switching templates. Text processing involves normalization, Unicode cleanup, number verbalization, partial diacritization (applied only to homographs), and Levantine lexicon overrides. Speech synthesis is performed using the Lahgtna-OmniVoice model, a zero-shot TTS model fine-tuned for Levantine Arabic dialect, with voice cloning from reference audio. The dataset features audio at a 24,000 Hz sampling rate, along with text transcripts, speaker IDs, speaker names, gender, and sentence type fields.




