naija_african_voices
收藏资源简介:
WaZoBiaSpeech 是一个大规模、高质量、完全转录的尼日利亚皮钦语(pcm)语音数据集。该数据集旨在加速非洲语境下的语音技术发展,促进语言多样性,并支持低资源机器学习研究。数据包含脚本化(scripted)和非脚本化(unscripted)录音,通过符合伦理的社区驱动流程收集,涵盖广泛的人口统计信息。数据集共包含 390,767 个音频片段,总时长约 1017 小时,来自 1119 位说话人。每个样本提供以下字段:音频唯一标识符(audio_id)、伪匿名说话人代码(speaker_id)、48 kHz 单声道 WAV 音频文件路径(audio_path)、人工精标转录文本(transcript)、音频时长(duration_seconds,单位秒)、性别(gender,男/女)、年龄段(age_group,15–29, 30–45, 46–60, 60+)、教育水平(education,初级/中级/高级)、内容领域(domain,如 EV, HE, AG, BU 等)、录音类型(type,脚本/非脚本)、数据划分(split,train/dev/dev_test,说话人无重叠)、语言代码(language,pcm)。数据集划分为训练集(335,288 个样本)、开发集(17,737 个样本)和开发测试集(17,324 个样本)。该数据集主要适用于自动语音识别(ASR)训练、低资源非洲语言自然语言处理、跨语言学习与迁移学习研究、多语言 ASR 系统评估以及语言学研究。为保护说话人隐私并防止语音滥用,数据集严格禁止用于语音克隆、声纹识别、监控或任何涉及识别或模仿个体的商业应用。数据集采用 Creative Commons Attribution 4.0 (CC BY 4.0) 许可证发布。
WaZoBiaSpeech is a large-scale, high-quality, fully transcribed Nigerian Pidgin (pcm) speech dataset. It aims to accelerate speech technology development in the African context, promote linguistic diversity, and support low-resource machine learning research. The data includes scripted and unscripted recordings, collected through an ethical community-driven process, covering a wide range of demographic information. The dataset contains a total of 390,767 audio clips, approximately 1,017 hours in duration, from 1,119 speakers. Each sample provides the following fields: audio unique identifier (audio_id), pseudonymized speaker code (speaker_id), 48 kHz mono WAV audio file path (audio_path), manually refined transcription text (transcript), audio duration in seconds (duration_seconds), gender (gender, male/female), age group (age_group, 15–29, 30–45, 46–60, 60+), education level (education, primary/secondary/tertiary), content domain (domain, e.g., EV, HE, AG, BU), recording type (type, scripted/unscripted), data split (split, train/dev/dev_test with non-overlapping speakers), and language code (language, pcm). The dataset is split into training set (335,288 samples), development set (17,737 samples), and development test set (17,324 samples). This dataset is mainly suitable for automatic speech recognition (ASR) training, low-resource African language NLP, cross-lingual and transfer learning research, multilingual ASR system evaluation, and linguistic research. To protect speaker privacy and prevent voice misuse, the dataset is strictly prohibited from being used for voice cloning, speaker recognition, surveillance, or any commercial application involving identification or impersonation of individuals. The dataset is released under the Creative Commons Attribution 4.0 (CC BY 4.0) license.
数据集概述
基本信息
- 数据集名称:WaZoBiaSpeech: 1,000+ Hour Nigerian Pidgin (pcm) Corpus
- 数据集地址:https://huggingface.co/datasets/Africanvoice/naija_african_voices
- 版本:30 Nov 2025
- 许可证:Creative Commons Attribution 4.0 (CC BY 4.0)
- 任务类别:自动语音识别(automatic-speech-recognition)
- 维护方:Data Science Nigeria / EqualyzAI
- 最后更新:2025年11月30日
数据集简介
WaZoBiaSpeech 是一个面向尼日利亚皮钦语(Nigerian Pidgin, pcm) 的大规模、高质量、全转写语音数据集,旨在推动非洲语境下语音技术的发展,促进语言多样性并支持低资源机器学习研究。数据包含脚本化(scripted) 和非脚本化(unscripted) 录音,通过符合伦理、以社区为中心的方式采集,具有广泛的人口统计覆盖。
语言覆盖
| 语言 | 总片段数 | 总时长 | 说话人数 |
|---|---|---|---|
| Nigerian Pidgin | 390,767 | 1017.04 h | 1119 |
数据集结构与特征
| 字段 | 类型 | 描述 |
|---|---|---|
audio_id |
string | 唯一音频标识符 |
speaker_id |
string | 匿名化说话人代码 |
audio_path |
Audio | 48 kHz 单声道 WAV 音频文件 |
transcript |
string | 干净的人工转写 |
duration_seconds |
float64 | 音频片段时长(秒) |
gender |
string | 性别(Male / Female) |
age_group |
string | 年龄段(15–29、30–45、46–60、60+) |
education |
string | 教育程度(Primary、Secondary、Tertiary) |
domain |
string | 内容语境(EV、HE、AG、BU) |
type |
string | 类型(Scripted / Unscripted) |
split |
string | 数据划分(train / dev / dev_test,按说话人分离) |
language |
string | 语言代码(如 pcm) |
数据划分
| 划分 | 字节数 | 样本数 |
|---|---|---|
| train | 419,304,709,289 | 335,288 |
| dev_test | 22,685,260,305 | 17,324 |
| dev | 19,624,690,744 | 17,737 |
- 下载大小:461,614,832,663 字节
- 数据集大小:461,614,660,338 字节
加载方式
推荐环境
bash pip install --upgrade datasets[audio] pip install --upgrade ffmpeg ffmpeg-python
标准加载
python from datasets import load_dataset
加载完整训练集
ds_train = load_dataset("Africanvoice/naija_african_voices", "default", split="train")
加载特定划分(如开发集)
ds_dev = load_dataset("Africanvoice/naija_african_voices", "default", split="dev")
流式模式(内存高效)
python from datasets import load_dataset
以流式模式加载 dev_test 划分
ds_stream = load_dataset( "Africanvoice/African_voices_yoruba", "default", split="dev_test", streaming=True )
预期用途与应用
- 自动语音识别(ASR)训练
- 低资源非洲语言的自然语言处理(NLP)
- 跨语言学习与迁移学习研究
- 多语言 ASR 系统评估
- 语言学研究和口音/方言建模
使用限制与局限
严格禁止的用途
- 语音克隆或适配(文本转语音/TTS)
- 声纹识别、说话人识别或模仿
- 监控、画像,或任何依赖识别或模仿个人的商业应用
局限
- 区域口音变体虽覆盖广泛,但并非完全详尽
- 自发性(非脚本化)语音片段可能包含自然的低水平背景噪声
- 不适用于生物识别或取证用途
引用信息
bibtex @dataset{wazobiaspeech-2025, title = {WaZoBiaSpeech: A 2,500-Hour Multilingual Speech Corpus for Hausa, Igbo, Nigerian Pidgin, and Yoruba}, author = {EqualyzAI and African Voices Team}, year = {2025}, url = {https://huggingface.co/datasets/Africanvoice/African_voices_yoruba}, note = {Niaja subset (pcm) version}, type = {dataset} }
联系方式
如有问题、疑问或合作意向,请在仓库中提交 issue 或直接联系维护者。





