Arabic Little STT
收藏资源简介:
Arabic Little STT数据集是一个包含288名6-13岁儿童在课堂环境中录制的355个黎凡特阿拉伯语语音的数据集。该数据集旨在填补阿拉伯语儿童语音数据稀缺的空白,并评估了最新的人工智能语音识别模型在儿童语音上的表现。数据集的创建过程中,所有音频均使用标准智能手机麦克风在课堂环境中录制,并经过人工转录和阿拉伯语特定规范化处理。该数据集可用于开发针对儿童语音的语音识别系统,并提高儿童在教育技术中的参与度。
The Arabic Little STT Dataset is a collection of 355 Levantine Arabic speech recordings collected from 288 children aged 6 to 13 in classroom environments. This dataset aims to fill the gap of scarce Arabic children's speech data, and evaluate the performance of state-of-the-art AI speech recognition models on children's speech. During the dataset construction, all audio recordings were captured using standard smartphone microphones in classroom settings, and underwent manual transcription and Arabic-specific normalization processing. This dataset can be used to develop speech recognition systems tailored for children's speech, and enhance children's engagement in educational technology.




