RobotsMali/afvoices
收藏资源简介:
--- dataset_info: - config_name: human-corrected features: - name: text dtype: string - name: duration dtype: float64 - name: audio dtype: audio - name: label-v1 dtype: string - name: label-v2 dtype: string splits: - name: train num_bytes: 62771143761 num_examples: 253290 - name: test num_bytes: 1515394591 num_examples: 6718 download_size: 59319505964 dataset_size: 64286538352 - config_name: model-annotated features: - name: duration dtype: float64 - name: audio dtype: audio - name: label-v1 dtype: string - name: label-v2 dtype: string splits: - name: train num_bytes: 55616591334 num_examples: 355571 download_size: 66321575877 dataset_size: 55616591334 - config_name: short features: - name: audio dtype: audio - name: duration dtype: float64 - name: label-v1 dtype: string - name: label-v2 dtype: string splits: - name: train num_bytes: 16345361845 num_examples: 259183 download_size: 16319374978 dataset_size: 16345361845 configs: - config_name: human-corrected data_files: - split: train path: human-corrected/train-* - split: test path: human-corrected/test-* - config_name: model-annotated data_files: - split: train path: model-annotated/train-* - config_name: short data_files: - split: train path: short/train-* license: cc-by-4.0 task_categories: - automatic-speech-recognition language: - bm tags: - bambara - African-Next-Voices - ANV - RobotsMali - afvoices - asr pretty_name: Robots --- # 📘 **African Next Voices – Bambara (AfVoices)** The **AfVoices** dataset is the largest open corpus of spontaneous Bambara speech at its release in late 2025. It contains **423 hours** of segmented audio and **612 hours** of original raw recordings collected across southern Mali. Speech was recorded in natural, conversational settings and annotated using a semi-automated transcription pipeline combining ASR pre-labels and human corrections. We release all the data processing code on [GitHub](https://github.com/RobotsMali-AI/afvoices). --- ## 🔎 **Quick Facts** | Category | Value | | ---------------------------------------- | ------------------------------------------------------------------------------------------------- | | **Total raw hours** | 612 h (1,777 raw recordings; publicly available on GCS) | | **Total segmented hours** | 423 h (874,762 segments) | | **Speakers** | 512 | | **Regions** | Bamako, Ségou, Sikasso, Bagineda, Bougouni | | **Avg. segment duration** | ~2 seconds | | **Subsets** | 159 h human-corrected, 212 h model-annotated, 52 h short (<1s) | | **Age distribution** | Broad, across young to elderly speakers (90% between 18 and 45) | | **Topics** | Health, agriculture, Miscellaneous (art, education, history etc.) | | **SNR distribution (raw recordings)** | 71.75% High or Very High SNR | | **Train / Test split** | 155 h / 4 h | --- ## **Motivation** The **African Next Voices (ANV)** project is a multi-country effort aiming to gather over **9,000 hours of speech** across 18 African languages. Its goal is to build high-quality datasets that empower local communities, support inclusive AI research, and provide strong foundations for ASR in underrepresented languages. As part of this initiative, **RobotsMali** led the Bambara data collection for Mali. This dataset reflects RobotsMali’s broader mission to advance AI and NLP research malian languages, with a long-term focus on improving education, access, and technology across Mali and the wider Manding linguistic region. --- ## 🎙️ **Characteristics of the Dataset** ### **Data Collection** * Speech was collected through trained **facilitators** who guided participants, ensured audio quality, and encouraged natural, topic-focused conversations. * All recordings are **spontaneous speech**, not read text. * A custom **Flutter mobile app** ([open-source](https://github.com/RobotsMali-AI/Africa-Voice-App)) was used to simplify the process and reduce training time. * Geographic focus: **Southern Mali**, to limit extreme accent variation and build a clean baseline corpus. ### **Segmentation and Preprocessing** * Raw audio was segmented using **Silero VAD**, retaining ~70% of the original duration. * Segments range from **240 ms to 30 s**. * Voice activity detection helped remove long silences and improve data usability. ### **Transcriptions** * Pre-transcribed using the ASR model **soloni-114m-tdt-ctc-v0**. * Human annotators corrected the transcripts. * A second model (**soloni-114m-tdt-ctc-v2**) was trained using the corrected transcripts and used to regenerate improved labels. * Two automatic transcription variants exist for each sample: **v1** (from soloni-v0) and **v2** (from soloni-v2). ### **Acoustic Event Tags** The following tags appear in transcriptions: | Tag | Meaning | | --------- | ------------------------------------------------------------- | | `[um]` | Vocalized pauses, filler sounds | | `[cs]` | Code-switched or foreign word | | `[noise]` | Background noise (applause, coughing, children, etc.) | | `[?]` | Inaudible or overlapped speech | | `[pause]` | Long silence (>5 seconds or >3 seconds at segment boundaries); due to VAD segmentation this tag is rarely used | --- ## 📂 **Subsets** ### **1. Human-corrected (159 h, 260k samples)** * Fully reviewed and corrected by annotators. * Only subset with a definitive `text` field containing the validated transcription. ### **2. Model-annotated (212 h, 355k samples)** * Includes automatic labels: `v1` (soloni-v0) and `v2` (soloni-v2). * No human review. ### **3. Short subset (52 h, 259k samples)** * Segments <1 second (formulaic expressions, discourse markers). * Excluded from human annotation for optimization purposes. * Automatically labeled (v1 & v2). --- ## ⚠️ **Limitations** * **Clean dataset vs real-world noise:** Over 70% of recordings can be categorized as relatively clean speech. Models trained solely on this dataset may underperform in noisy street or radio environments typical in Mali. See this [report](https://zenodo.org/records/17672774) if you are interested in learning more about the strengths and weaknesses of RobotsMali's ASR models. * **Reduced code-switching:** French terms were often replaced by `[cs]` or normalized into Bambara phonology. This improves model stability but reduces realism for natural bilingual speech. * **Geographic homogeneity:** Focused on the southern region to control accent variability. Broader dialectal coverage might require additional data. * **Simplified linguistic conditions:** Overlaps, multi-speaker settings, and conversational chaos are minimized—again improving training stability at the cost of deployment realism. --- ## 📑 **Citation** ```bibtex @misc{diarra2025dealinghardfactslowresource, title={Dealing with the Hard Facts of Low-Resource African NLP}, author={Yacouba Diarra and Nouhoum Souleymane Coulibaly and Panga Azazia Kamaté and Madani Amadou Tall and Emmanuel Élisé Koné and Aymane Dembélé and Michael Leventhal}, year={2025}, eprint={2511.18557}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2511.18557}, } ``` --- You may want to download the original 612 hours dataset with its associated metadata for research purposes or to create a derivative. You will find the codes and manifest files to download those files from Google Cloud Storage in this repository: [RobotsMali-AI/afvoices](https://github.com/RobotsMali-AI/afvoices). Do not hesitate to open an issue for Help or suggestions 🤗
dataset_info: - 配置名称:human-corrected 特征: - 名称:text 数据类型:字符串(string) - 名称:duration 数据类型:64位浮点数(float64) - 名称:audio 数据类型:音频(audio) - 名称:label-v1 数据类型:字符串(string) - 名称:label-v2 数据类型:字符串(string) 划分集: - 名称:train 字节数:62771143761 样本数:253290 - 名称:test 字节数:1515394591 样本数:6718 下载大小:59319505964 数据集总大小:64286538352 - 配置名称:model-annotated 特征: - 名称:duration 数据类型:64位浮点数(float64) - 名称:audio 数据类型:音频(audio) - 名称:label-v1 数据类型:字符串(string) - 名称:label-v2 数据类型:字符串(string) 划分集: - 名称:train 字节数:55616591334 样本数:355571 下载大小:66321575877 数据集总大小:55616591334 - 配置名称:short 特征: - 名称:audio 数据类型:音频(audio) - 名称:duration 数据类型:64位浮点数(float64) - 名称:label-v1 数据类型:字符串(string) - 名称:label-v2 数据类型:字符串(string) 划分集: - 名称:train 字节数:16345361845 样本数:259183 下载大小:16319374978 数据集总大小:16345361845 配置项: - 配置名称:human-corrected 数据文件: - 划分集:train 路径:human-corrected/train-* - 划分集:test 路径:human-corrected/test-* - 配置名称:model-annotated 数据文件: - 划分集:train 路径:model-annotated/train-* - 配置名称:short 数据文件: - 划分集:train 路径:short/train-* 许可证:cc-by-4.0 任务类别: - 自动语音识别(automatic-speech-recognition) 语言: - 班巴拉语(bm) 标签: - bambara(班巴拉语) - African-Next-Voices - ANV - RobotsMali - afvoices - asr(自动语音识别) 展示名称:Robots # 📘 **非洲之声未来计划——班巴拉语(AfVoices)** **AfVoices** 数据集是2025年末发布时规模最大的开源班巴拉语自发语音语料库。其包含在马里南部采集的423小时分段音频与612小时原始未处理录音。语音数据采用自然会话场景录制,并通过结合自动语音识别(ASR)预标注与人工校正的半自动化转录流程完成标注。我们已将所有数据处理代码开源至 "[GitHub](https://github.com/RobotsMali-AI/afvoices)"。 --- ## 🔎 **核心概况** | 类别 | 数值 | | --- | --- | | **总原始录音时长** | 612小时(共1777条原始录音,可在谷歌云存储(GCS)公开获取) | | **总分段音频时长** | 423小时(共874,762个音频分段) | | **说话者数量** | 512位 | | **采集区域** | 巴马科、塞古、锡卡索、巴吉内达、布古尼 | | **平均分段时长** | 约2秒 | | **子集划分** | 159小时人工校正子集、212小时模型标注子集、52小时短分段子集(时长<1秒) | | **年龄分布** | 覆盖青年至老年群体,其中90%的说话者年龄介于18至45岁 | | **话题范畴** | 健康、农业、其他杂项(艺术、教育、历史等) | | **原始录音信噪比(SNR)分布** | 71.75%为高或极高信噪比 | | **训练集/测试集划分** | 155小时 / 4小时 | --- ## **项目初衷** 非洲之声未来计划(African Next Voices,简称ANV)是一项多国合作项目,目标是采集覆盖18种非洲语言的总计9000余小时语音数据。该项目旨在构建高质量数据集,赋能本地社区发展,支持包容性人工智能(AI)研究,并为弱势语言的自动语音识别(ASR)研究奠定坚实基础。 作为该计划的一部分,RobotsMali团队负责牵头马里地区的班巴拉语数据采集工作。本数据集体现了RobotsMali推动马里及更广曼丁语族地区人工智能与自然语言处理(NLP)研究的整体使命,长期目标是改善马里及周边地区的教育、信息获取与技术发展水平。 --- ## 🎙️ **数据集特性** ### **数据采集** * 语音数据由经过培训的**协调员**采集,他们负责引导参与者、保障录音质量并鼓励开展自然的主题式对话。 * 所有录音均为**自发口语**,而非朗读文本。 * 项目采用定制化**Flutter移动应用**("[开源地址](https://github.com/RobotsMali-AI/Africa-Voice-App)")简化流程并缩短训练周期。 * 采集区域聚焦**马里南部**,以限制口音差异过大的问题,构建干净的基准语料库。 ### **分段与预处理** * 原始音频通过**Silero VAD**进行分段,保留了约70%的原始录音时长。 * 音频分段时长范围为240毫秒至30秒。 * 语音活动检测(VAD)可移除长静音片段,提升数据可用性。 ### **转录标注** * 首先使用自动语音识别模型**soloni-114m-tdt-ctc-v0**生成预转录结果。 * 随后由人工标注员对转录结果进行校正。 * 研究人员利用校正后的转录结果训练了第二个模型(**soloni-114m-tdt-ctc-v2**),并使用该模型重新生成优化后的标注。 * 每个样本均包含两种自动转录变体:**v1**(基于soloni-v0模型生成)与**v2**(基于soloni-v2模型生成)。 ### **声学事件标签** 转录文本中包含以下标签: | 标签 | 含义 | | --- | --- | | `[um]` | 带音停顿、填充语 | | `[cs]` | 语码转换或外来词 | | `[noise]` | 背景噪音(如掌声、咳嗽、儿童声响等) | | `[?]` | 无法听清或重叠的语音 | | `[pause]` | 长静音(>5秒或分段边界处>3秒);由于VAD分段的特性,该标签极少使用 | --- ## 📂 **子集划分** ### **1. 人工校正子集(159小时,26万条样本)** * 经标注员全面审核与校正。 * 是唯一包含明确`text`字段(存储经过验证的转录文本)的子集。 ### **2. 模型标注子集(212小时,35.5万条样本)** * 包含自动生成的标注:`v1`(基于soloni-v0)与`v2`(基于soloni-v2)。 * 未经过人工审核。 ### **3. 短分段子集(52小时,25.9万条样本)** * 分段时长<1秒(多为固定表达、话语标记语)。 * 出于优化考虑,未对其进行人工标注。 * 仅通过自动方式生成标注(v1与v2)。 --- ## ⚠️ **局限性说明** * **干净数据集与真实场景噪音:** 超过70%的录音可归类为相对干净的语音。仅基于本数据集训练的模型,在马里常见的嘈杂街道或广播场景中可能表现不佳。若需了解RobotsMali自动语音识别模型的优缺点详情,可参阅该"[报告](https://zenodo.org/records/17672774)"。 * **语码转换受限:** 法语词汇常被替换为`[cs]`标签或被标准化为班巴拉语语音形式。此举提升了模型稳定性,但降低了自然双语口语的真实感。 * **地域同质化:** 采集区域聚焦南部以控制口音差异。若需覆盖更广的方言范围,则需补充额外数据。 * **语言场景简化:** 重叠语音、多说话者场景以及会话混乱情况均被最小化——此举同样提升了训练稳定性,但牺牲了部署场景的真实度。 --- ## 📑 **引用格式** bibtex @misc{diarra2025dealinghardfactslowresource, title={应对低资源非洲自然语言处理的现实挑战}, author={雅库巴·迪亚拉、努胡姆·苏莱曼·库利巴利、潘加·阿扎齐亚·卡马泰、马达尼·阿马杜·塔尔、埃马纽埃尔·埃利泽·科内、艾曼内·登贝莱、迈克尔·莱文索尔}, year={2025}, eprint={2511.18557}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2511.18557}, } --- 若您出于研究目的或衍生开发需求,可下载包含元数据的612小时原始数据集。您可在本仓库 "[RobotsMali-AI/afvoices](https://github.com/RobotsMali-AI/afvoices)" 中找到用于从谷歌云存储下载相关文件的代码与清单文件。如有任何疑问或建议,欢迎提交Issue 🤗



