ghananlpcommunity/ghana-named-entities-tts-twi
收藏资源简介:
Ghana Named Entities TTS — Twi 是一个Twi语(加纳阿坎语方言)的语音合成数据集,内容基于加纳命名实体(人物、地点、组织、概念)的描述文本。每个音频片段是描述多个命名实体的段落的合成朗读。数据集包含6,704个样本,音频格式为24kHz单声道WAV,采用女性青年成人中等音高的合成语音生成。数据集构建流程包括四个步骤:首先从GhanaNLP社区收集约253,000个加纳命名实体及其英文描述;然后将英文描述按每组10个实体分块组合成段落;接着使用Gemini模型将英文段落翻译成Twi语;最后使用OmniVoice模型将Twi语段落合成为语音。该数据集旨在用于训练和评估Twi语TTS和ASR系统、低资源非洲语言语音研究以及加纳专有名词的发音建模。需要注意的是,音频为合成语音,可能存在韵律和发音方面的TTS伪影;每个片段覆盖多个实体组成的段落而非孤立话语;翻译为自动化处理,部分Twi语表达可能不够地道。
Ghana Named Entities TTS — Twi is a Twi-language speech dataset built from descriptions of Ghana named entities (people, places, organisations, and concepts). Each audio clip is a synthesised reading of a passage that describes several named entities. The dataset contains 6,704 examples in WAV format at 24,000 Hz mono, generated with a synthetic female young adult voice at moderate pitch. The dataset was created through a four-step pipeline: source collection of ~253,000 Ghana named entities with English descriptions from GhanaNLP Community; chunking English descriptions into passages of 10 entities each; translation of passages into Twi using Gemini model; and speech synthesis of Twi passages using OmniVoice model. The dataset is intended for training and evaluating Twi TTS and ASR systems, low-resource African-language speech research, and named-entity pronunciation modelling for Ghanaian proper nouns. Limitations include: audio is synthetic and may contain TTS artefacts; each clip covers multiple entities concatenated into a passage; translation was automated and some Twi renderings may not be idiomatic.




