THCHS-30
收藏资源简介:
THCHS-30是由清华大学信息技术研究院语音与语言技术研究中心发布的免费中文语音数据库,旨在支持语音识别研究。该数据集包含超过30小时的普通话语音信号,由50名参与者录制,采样率为16,000 Hz,样本大小为16位。数据集内容丰富,包括1000个从新闻中选取的句子,旨在增强863数据库的语音覆盖。此外,数据集还提供了包括词汇、语言模型和训练配方在内的全套资源,支持构建大型词汇连续中文语音识别系统。THCHS-30的应用领域广泛,主要用于解决中文语音识别中的数据获取难题,尤其适合初入此领域的年轻研究者。
THCHS-30 is a free Mandarin Chinese speech database released by the Center for Speech and Language Technology, Research Institute of Information Technology, Tsinghua University, aimed at supporting speech recognition research. This dataset contains over 30 hours of Mandarin speech signals recorded by 50 participants, with a sampling rate of 16,000 Hz and a 16-bit sample depth. It features rich content, including 1,000 sentences selected from news sources, with the purpose of enhancing the speech coverage of the 863 Database. In addition, the dataset provides a complete set of resources including lexicons, language models, and training recipes, which supports the development of large-vocabulary continuous Mandarin speech recognition systems. THCHS-30 has a wide range of application scenarios, mainly used to address the data acquisition challenges in Chinese speech recognition, and is particularly suitable for young researchers new to this field.




