Xasan-Speech-Phoneme-coverage
收藏资源简介:
Xasan Speech Phoneme Coverage 是一个索马里语语音数据集,专为文本到语音(TTS)模型的微调而设计。该数据集包含 1,251 个配对的音频和文本样本,音频为 WAV 格式,采样率为 24 kHz,所有录音由单一男性说话人(Xasan)在受控环境下录制,确保音频干净、高质量且无背景噪音。数据集的亮点在于实现了 100% 的音素覆盖率,每个句子都是唯一的,无重复内容,覆盖了除 P 和 V 外的所有本地索马里字母和音素(P 和 V 非索马里语原生字母)。数据集格式简单,每个样本包含 audio 和 text 两个字段,适用于微调现有的 TTS 模型,但由于规模较小(1k<n<10k),不推荐用于从头训练基础模型。该数据集针对低资源语言(索马里语作为非洲语言)场景,支持语音合成和语音相关任务,遵循 Apache-2.0 许可证。
Xasan Speech Phoneme Coverage is a Somali speech dataset designed for fine-tuning text-to-speech (TTS) models. It contains 1,251 paired audio and text samples, with audio in WAV format at a 24 kHz sampling rate. All recordings are made by a single male speaker (Xasan) in a controlled environment, ensuring clean, high-quality audio with no background noise. The datasets highlight is 100% phoneme coverage, with each sentence being unique and non-repetitive, covering all native Somali letters and phonemes except P and V (which are non-native Somali letters). The dataset format is simple, with each sample containing audio and text fields, making it suitable for fine-tuning existing TTS models, but not recommended for training base models from scratch due to its small size (1k<n<10k). It targets low-resource language scenarios (Somali as an African language), supports speech synthesis and related tasks, and follows the Apache-2.0 license.




