THCHS-30
收藏资源简介:
“THCHS30是由清华大学语音与语言技术中心(CSLT)发布的开放式汉语语音数据库。原始录音是2002年在清华大学国家重点实验室的朱晓燕教授的指导下,由王东完成的。清华大学计算机科学系智能与系统,原名“TCMSD”,意思是“清华连续普通话语音数据库”,时隔13年出版,由王东博士发起,并得到了教授的支持。朱小燕。我们希望为语音识别领域的新研究人员提供一个玩具数据库。因此,该数据库对学术用户完全免费。整个软件包包含建立中文语音识别所需的全套语音和语言资源系统。”
THCHS30 is an open-source Mandarin speech database released by the Center for Speech and Language Technology (CSLT) at Tsinghua University. The original recordings were completed by Wang Dong in 2002 under the guidance of Professor Zhu Xiaoyan at the State Key Laboratory of Tsinghua University. Originally developed by the Intelligence and Systems Group of the Department of Computer Science and Technology at Tsinghua University, this resource was initially named "TCMSD", which stands for "Tsinghua Continuous Mandarin Speech Database". It was republished 13 years later, initiated by Dr. Wang Dong and supported by Professor Zhu Xiaoyan. We aim to provide a beginner-friendly toy database for new researchers in the field of speech recognition, so this database is completely free for academic users. The entire software package includes a full set of speech and linguistic resources required to build Chinese speech recognition systems.

- THCHS-30数据集首次发表,由清华大学语音与语言技术中心发布,旨在为中文语音识别研究提供一个标准化的数据集。
- THCHS-30数据集首次应用于多个中文语音识别研究项目,显著提升了模型的训练效果和识别准确率。
- THCHS-30数据集被广泛应用于学术界和工业界,成为中文语音识别领域的重要基准数据集之一。
- THCHS-30数据集的扩展版本发布,增加了更多的语音样本和多样化的语音场景,进一步丰富了数据集的内容和应用范围。



