SimbaBench_dataset
收藏资源简介:
SimbaBench 是一个面向非洲语言的多语言基准数据集,专注于自动语音识别(ASR)、文本到语音(TTS)和口语语言识别(SLID)任务。数据集包含多种语言的配置,每种配置提供了标准化的评估分割,涵盖样本数量和音频时长。数据集支持的语言包括南非荷兰语、阿姆哈拉语、班萨语等60多种非洲语言。数据集采用CC-BY-4.0许可,适用于低资源、多语言场景的研究和开发。示例代码展示了如何加载数据集进行模型评估。
SimbaBench is a multilingual benchmark dataset tailored for African languages, focusing on three core tasks: Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Spoken Language Identification (SLID). The dataset provides language-specific configurations, each paired with standardized evaluation splits covering both sample counts and total audio duration. It supports over 60 African languages including Afrikaans, Amharic, Basa, and more. Released under the CC-BY-4.0 license, the dataset is suitable for research and development in low-resource and multilingual scenarios. Example code is provided to demonstrate how to load the dataset for model evaluation.




