遇见数据集

BembaSpeech

收藏
arXiv2021-02-09 更新2024-06-21 收录
官方服务:

资源简介:

BembaSpeech是一个专为Bemba语言设计的自动语音识别数据集,由赞比亚大学计算机科学系创建。该数据集包含超过24小时的Bemba语言朗读语音,总计14,438条录音,适用于训练和测试自动语音识别系统。数据集内容来源于Bemba文学和其他在线资源,通过Lig-Aikuma应用收集。BembaSpeech旨在解决非洲语言资源稀缺的问题,特别是针对赞比亚的Bemba语言,支持构建更有效的语音识别技术。

BembaSpeech is an automatic speech recognition (ASR) dataset specifically designed for the Bemba language, created by the Department of Computer Science at the University of Zambia. This dataset contains over 24 hours of read speech in the Bemba language, with a total of 14,438 audio recordings, and is applicable for training and testing automatic speech recognition systems. The content of the dataset is sourced from Bemba literature and other online resources, and was collected via the Lig-Aikuma application. BembaSpeech aims to address the problem of scarce language resources for African languages, particularly the Bemba language in Zambia, to support the development of more effective speech recognition technologies.

创建时间:
2021-02-09
二维码
社区交流群
二维码
科研交流群
商业服务