uzbek-speech-corpus
收藏资源简介:
乌兹别克语语音语料库(USC)是由ISSAI和塔什干信息技术大学计算机系统系图像与语音处理实验室合作开发的。该语料库包含958个不同说话者的105小时转录音频。为确保高质量,语料库经过母语者的手动检查。USC主要用于自动语音识别(ASR),但也适用于语音合成和语音翻译等其他任务。据我们所知,USC是第一个在Creative Commons Attribution 4.0 International许可下开放给学术和商业用途的乌兹别克语语音语料库。我们期望USC将成为乌兹别克语ASR研究的基准数据集,并为语音研究社区提供宝贵的资源。
The Uzbek Speech Corpus (USC) was developed jointly by ISSAI and the Image and Speech Processing Laboratory of the Department of Computer Systems, Tashkent University of Information Technologies. This corpus contains 105 hours of transcribed audio from 958 distinct speakers. To ensure high quality, the corpus has been manually verified by native speakers. USC is primarily intended for automatic speech recognition (ASR), but also supports other tasks such as speech synthesis and speech translation. To the best of our knowledge, USC is the first Uzbek speech corpus made available for both academic and commercial use under the Creative Commons Attribution 4.0 International license. We anticipate that USC will serve as a benchmark dataset for Uzbek ASR research and provide a valuable resource for the speech research community.




