DiDiSpeech
收藏资源简介:
DiDiSpeech是一个大规模的普通话语音数据集,由滴滴出行创建,包含约800小时的48kHz采样率语音数据,来自6000名不同年龄和性别的母语普通话发音人。数据集在安静环境中录制,适用于多种语音处理任务,如语音转换、多说话人文本到语音合成和自动语音识别。创建过程包括使用移动设备在安静环境中录制,并经过信号处理和文本预处理。该数据集旨在支持学术研究和工业应用,特别是在提升普通话语音处理技术方面。
DiDiSpeech is a large-scale Mandarin speech dataset developed by DiDi Chuxing. It contains approximately 800 hours of speech data sampled at 48 kHz, collected from 6,000 native Mandarin speakers spanning various age groups and genders. Recorded in quiet environments, the dataset supports a broad spectrum of speech processing tasks, such as voice conversion, multi-speaker text-to-speech synthesis, and automatic speech recognition. The dataset's construction involves capturing audio via mobile devices in quiet settings, followed by signal processing and text preprocessing. This resource is designed to facilitate academic research and industrial applications, with a particular focus on advancing Mandarin speech processing technologies.




