african-languages-speech
收藏资源简介:
该数据集是WAXAL非洲语言语音数据集的过滤版本(第一波),包含豪萨语(hau_tts)、约鲁巴语(yor_tts)和伊博语(ibo_tts)三种语言的音频与转录配对样本。数据以tar归档文件形式组织,每个归档包含一个音频文件和一个JSON记录,JSON记录中包含转录文本、语言标签、说话人标识、来源ID和来源信息。数据集来源于Google的WaxalNLP数据集,经过流式逐行处理,移除了空文本、明显垃圾文本以及缺失音频的样本,确保了转录与音频的对齐,并使用随机水塘采样进行保留。每个配置均包含一个audit.json文件,用于记录数量统计、丢弃原因、校验和以及样本信息。该数据集适用于自动语音识别(ASR)和文本转语音(TTS)等语音任务,遵循CC-BY-4.0许可证。
This dataset is a filtered version (first wave) of the WAXAL African Language Speech Dataset, containing audio-transcription paired samples for three languages: Hausa (hau_tts), Yoruba (yor_tts), and Igbo (ibo_tts). The data is organized as tar archive files, each containing an audio file and a JSON record. The JSON record includes the transcription text, language tag, speaker identifier, source ID, and source information. The dataset originates from Googles WaxalNLP dataset and has been processed via streaming, removing empty texts, obvious garbage texts, and samples with missing audio, ensuring alignment between transcription and audio, and retained using random reservoir sampling. Each configuration includes an audit.json file that records quantity statistics, discard reasons, checksums, and sample information. This dataset is suitable for speech tasks such as automatic speech recognition (ASR) and text-to-speech (TTS), and is licensed under CC-BY-4.0.
数据集概述:WAXAL 非洲语言语音(过滤后第一批)
基本信息
- 数据集名称:WAXAL African-language speech — filtered wave 1
- 许可证:CC-BY-4.0
- 发布机构:VelkroLM
内容简介
该数据集包含经过筛选和限制的 WAXAL 音频与转录文本配对数据,涵盖三种非洲语言:
- 豪萨语(Hausa):配置标识为
hau_tts - 约鲁巴语(Yoruba):配置标识为
yor_tts - 伊博语(Igbo):配置标识为
ibo_tts
每个语言的配置均被分割为多个 tar 压缩包,每个压缩包包含一个音频文件及对应的 JSON 记录。JSON 记录中收录了转录文本、语言、说话人、来源 ID 和数据出处信息。
数据来源
- 上游数据集:google/WaxalNLP
- 相关论文:arXiv:2602.02734
- 该数据集为上游数据的衍生版本,不授予超出上游条款的任何权利。用户需查阅上游数据集卡片以了解具体的组件许可和署名要求。
处理流程
数据收集者采用逐行流式处理方式,主要步骤包括:
- 删除空文本、明显的垃圾信息文本及缺失的音频
- 保持转录文本与音频的对齐
- 使用带种子的随机水库采样方法进行抽样
- 在成功验证后删除原始暂存数据
每个配置均包含一份 audit.json 审计文件,记录数据计数、丢弃原因、校验和及样本信息。
后续计划
其余的 WAXAL 配置将按分批节奏逐步添加,而不会复制到第二个本地副本中。




