Mandarin Dysarthria Speech Corpus (MDSC)
收藏资源简介:
Mandarin Dysarthria Speech Corpus (MDSC)是由北京AISHELL科技有限公司与中国科学技术大学合作创建的,专为家庭环境中语音障碍患者设计的语音数据集。该数据集包含18,630条录音,总计17小时,其中9.4小时来自21位语音障碍患者,7.6小时来自25位标准语音模式的演讲者。数据集涵盖了年龄、性别、疾病类型和可理解性评估等信息,旨在通过提供多样化的语音样本,解决语音唤醒技术在语音障碍患者中的应用问题。创建过程中,录音包括10个唤醒词和355个非唤醒词,采用16kHz采样率,在安静的室内环境中进行。MDSC的应用领域主要集中在改善语音障碍患者的语音唤醒系统,提高其生活质量。
Mandarin Dysarthria Speech Corpus (MDSC) was co-developed by Beijing AISHELL Technology Co., Ltd. and the University of Science and Technology of China. It is a speech corpus specifically designed for dysarthric patients in home environments, containing 18,630 audio recordings totaling 17 hours: 9.4 hours from 21 dysarthric patients, and the remaining 7.6 hours from 25 speakers with standard speech patterns. The corpus covers comprehensive metadata including age, gender, disease type, and intelligibility assessment results, aiming to address the application barriers of speech wake-up technology for dysarthric patients by providing diverse speech samples. During data collection, the recordings covered 10 wake-up words and 355 non-wake-up words, with a sampling rate of 16 kHz, and all recordings were conducted in quiet indoor environments. The primary application scope of MDSC is to optimize speech wake-up systems for dysarthric patients, thereby improving their quality of life.




