Multi-Domain Cantonese Corpus (MDCC)
收藏资源简介:
Multi-Domain Cantonese Corpus (MDCC)是由香港科技大学创建的一个包含73.6小时干净朗读语音的数据集,涵盖哲学、政治、教育、文化、生活方式和家庭等多个领域。该数据集从香港的粤语有声读物中收集,包含约83,275条语音记录,每条记录时长在0.22至15秒之间。MDCC旨在解决粤语自动语音识别(ASR)系统中数据稀缺的问题,并通过与现有数据集如Common Voice zh-HK的比较,展示了其在ASR研究中的有效性。
The Multi-Domain Cantonese Corpus (MDCC) is a dataset containing 73.6 hours of clean read speech, developed by The Hong Kong University of Science and Technology. It covers multiple domains including philosophy, politics, education, culture, lifestyle, and family. Collected from Cantonese audiobooks in Hong Kong, this dataset comprises approximately 83,275 speech records, with each clip ranging from 0.22 to 15 seconds in duration. MDCC aims to address the issue of data scarcity in Cantonese automatic speech recognition (ASR) systems, and demonstrates its effectiveness in ASR research through comparisons with existing datasets such as Common Voice zh-HK.
- 1Automatic Speech Recognition Datasets in Cantonese: A Survey and New Dataset香港科技大学 · 2022年



