AISHELL-3
收藏资源简介:
AISHELL-3是由北京壳牌科技有限公司创建的大型高质量多说话人普通话语音数据集,旨在训练多说话人文本到语音(TTS)系统。该数据集包含约85小时的中性情感录音,由218位母语为普通话的中国说话人录制,涵盖性别、年龄组和方言等辅助属性。数据集提供汉字级和拼音级转录,适用于构建能够实现零样本语音克隆的鲁棒合成模型。AISHELL-3的应用领域包括智能语音命令、新闻报道和地理信息等,旨在解决普通话TTS系统训练数据不足的问题。
AISHELL-3 is a large-scale, high-quality multi-speaker Mandarin speech dataset developed by Beijing Shell Technology Co., Ltd., which is purpose-built for training multi-speaker text-to-speech (TTS) systems. The dataset encompasses approximately 85 hours of neutral-emotion audio recordings, collected from 218 native Mandarin speakers across China, with auxiliary metadata covering gender, age groups, dialects and other related attributes. It provides both character-level and pinyin-level transcriptions, making it suitable for constructing robust speech synthesis models that support zero-shot voice cloning. The application scenarios of AISHELL-3 include intelligent voice commands, news reporting, geographic information services and more, and it aims to address the shortage of training data for Mandarin TTS systems.

- 1AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines武汉大学计算机科学学院 · 2021年



