WaxalNLP
收藏资源简介:
Waxal项目为非洲语言提供了自动语音识别(ASR)和文本转语音(TTS)数据集。创建和发布这些数据集的目的是促进研究,提高这些服务不足语言的语音和语言技术的准确性和流畅性,并作为数字保存的存储库。Waxal数据集是通过与马凯雷雷大学、加纳大学、Digital Umuganda和Media Trust的合作收集的,由谷歌和盖茨基金会资助,协议要求数据集公开可访问。ASR数据集包含14种非洲语言的大约1,250小时的转录自然语音,代表超过1亿使用者的40个撒哈拉以南非洲国家。TTS数据集包含10种非洲语言的大约240小时的脚本自然语音。
The Waxal Project curates automatic speech recognition (ASR) and text-to-speech (TTS) datasets tailored for African languages. These datasets were developed and released to advance academic research, improve the accuracy and fluency of speech and language technologies for these under-resourced languages, and serve as a dedicated repository for digital preservation efforts. The Waxal datasets were collected through partnerships with Makerere University, University of Ghana, Digital Umuganda, and Media Trust, with funding provided by Google and the Bill & Melinda Gates Foundation. Per the terms of the funding agreement, the datasets must be publicly accessible. The ASR dataset comprises approximately 1,250 hours of transcribed natural speech spanning 14 African languages, which are spoken across 40 Sub-Saharan African countries by a combined speaker population of over 100 million. The TTS dataset contains approximately 240 hours of scripted natural speech across 10 African languages.
Waxal NLP 数据集概述
数据集基本信息
- 数据集名称: Waxal NLP Datasets
- 提供方: Google Research
- 许可证: CC-BY-SA-4.0, CC-BY-4.0 (具体语言许可证因数据提供方而异)
- 当前版本: 1.0.0
- 最后更新: 2026年1月
- 数据集地址: https://huggingface.co/datasets/google/WaxalNLP
数据集描述
Waxal项目提供用于非洲语言的自动语音识别和文本到语音数据集。其创建和发布的目标是促进研究,以提高这些服务不足语言的语音和语言技术的准确性和流畅性,并作为数字保存的存储库。
该数据集通过与马凯雷雷大学、加纳大学、Digital Umuganda和Media Trust的合作获取,由谷歌和盖茨基金会资助,并同意使数据集可公开访问。
任务与配置
- 任务类别: 自动语音识别, 文本到语音
- 配置:
asr: 自动语音识别数据tts: 文本到语音数据
语言覆盖
- 总语言数: 20种非洲语言
- 语言列表: Acholi, Akan, Dagbani, Dagaare, Ewe, Fante, Fula, Hausa, Igbo, Ikposo, Lingala, Luganda, Masaaba, Malagasy, Nyankole, Shona, Soga, Kiswahili, Twi, Yoruba
ASR数据集语言详情
- 语言数量: 14种
- 总时长: 约1250小时的自然语音转录数据
- 覆盖人群: 代表超过40个撒哈拉以南非洲国家的1亿多使用者
- 数据提供方与语言:
- 马凯雷雷大学: Acholi, Luganda, Masaaba, Nyankole, Soga (许可证: CC-BY-4.0)
- 加纳大学: Akan, Ewe, Dagbani, Dagaare, Ikposo (许可证: CC-BY-NC-4.0)
- Digital Umuganda: Fula, Lingala, Shona, Malagasy (许可证: CC-BY-4.0)
TTS数据集语言详情
- 语言数量: 10种
- 总时长: 约240小时的脚本自然语音数据
- 数据提供方与语言:
- 马凯雷雷大学: Acholi, Luganda, Kiswahili, Nyankole (许可证: CC-BY-4.0)
- 加纳大学: Akan (Fante, Twi) (许可证: CC-BY-NC-4.0)
- Media Trust: Fula, Igbo, Hausa, Yoruba (许可证: CC-BY-4.0)
数据集结构
数据字段
ASR配置字段:
id: 唯一标识符speaker_id: 说话者唯一标识符audio: 音频数据transcription: 音频转录文本language: ISO 639-2语言代码gender: 说话者性别 (Male, Female, 或空)
TTS配置字段:
id: 唯一标识符speaker_id: 说话者唯一标识符audio: 音频数据transcription: 转录文本locale: ISO 639-2语言代码gender: 说话者性别
数据划分
ASR数据集划分:
train: 80%的标注数据validation: 10%的标注数据test: 10%的标注数据unlabeled: 所有没有对应转录的样本
TTS数据集划分:
数据被划分为train、validation和test集,结构类似。
数据来源
- 加纳大学: UGSpeechData (https://doi.org/10.57760/sciencedb.22298)
- Digital Umuganda: AfriVoice (DigitalUmuganda/AfriVoice)
- 马凯雷雷大学: Yogera Dataset (https://doi.org/10.7910/DVN/BEROE0)
- Media Trust
使用注意事项
使用数据前请检查您所用具体语言的许可证,因为它们可能因提供方而异。




