遇见数据集

Abjad-Kids

收藏
Hugging Face2026-03-14 更新2026-03-16 收录
官方服务:

资源简介:

Abjad-Kids 是一个为初级教育应用设计的阿拉伯语语音分类数据集。它包含多个儿童说话者录制的阿拉伯字母、数字和颜色的发音,支持阿拉伯语儿童语音的自动语音识别、音频分类和教育技术研究。数据集分为三个主要类别:字母、数字和颜色,每个类别都有对应的子类别,如字母类别包含 Alam (أ)、Ba (ب)、Ta (ت) 等。数据集结构清晰,每个类别都有对应的 CSV 文件,包含音频文件路径和标签两列。适用于自动语音识别 (ASR)、音频分类任务、初级教育和语言学习应用以及发音评估和计算机辅助发音教学 (CAPT)。数据集采用 MIT 许可证。

Abjad-Kids is an Arabic speech classification dataset designed for primary education applications. It contains pronunciations of Arabic letters, numbers and colors recorded by multiple child speakers, supporting research on automatic speech recognition (ASR), audio classification and educational technology for Arabic child speech. The dataset is divided into three main categories: letters, numbers and colors, each with corresponding subcategories. For example, the letter category includes Alam (أ), Ba (ب), Ta (ت), and so on. The dataset has a clear structure, with corresponding CSV files for each category that contain two columns: audio file path and label. It is applicable to automatic speech recognition (ASR), audio classification tasks, primary education and language learning applications, as well as pronunciation assessment and computer-assisted pronunciation teaching (CAPT). The dataset is released under the MIT License.

创建时间:
2026-03-11
二维码
社区交流群
二维码
科研交流群
商业服务