CUCHILD
收藏资源简介:
CUCHILD是由香港中文大学电子工程系和耳鼻喉头颈外科系合作开发的大型粤语儿童语音语料库。该数据集包含1986名3至6岁儿童的口语词汇,涵盖正常发育儿童和语音障碍儿童。数据集设计用于支持科学和临床研究,以及与儿童语音评估相关的技术开发。数据集包括130个1至4音节的粤语词汇,覆盖所有粤语音素,旨在解决儿童语音识别、语音错误检测和说话人分割等问题。
CUCHILD is a large-scale Cantonese children’s speech corpus co-developed by the Department of Electronic Engineering and the Department of Otorhinolaryngology, Head and Neck Surgery of The Chinese University of Hong Kong. This dataset encompasses spoken vocabulary from 1,986 children aged 3 to 6 years, including both typically developing children and those with speech disorders. It is designed to facilitate scientific and clinical research, as well as the development of technologies pertaining to children’s speech assessment. The corpus comprises 130 Cantonese vocabulary items with 1 to 4 syllables, covering all Cantonese phonemes, and is intended to address core tasks including children’s speech recognition, speech error detection, and speaker diarization.

- 1CUCHILD: A Large-Scale Cantonese Corpus of Child Speech for Phonology and Articulation Assessment香港中文大学 · 2020年



