TextrolSpeech
收藏资源简介:
TextrolSpeech是由浙江大学创建的大型开放源代码语音情感数据集,包含330小时的语音数据和236,220对自然文本风格提示。该数据集通过多阶段提示编程方法生成,涵盖五种风格因素,包括性别、音高、语速、音量和情感,旨在推动文本可控TTS系统的发展。数据集的创建过程中,整合了LibriTTS和VCTK数据集,并额外收集了多个情感数据集以丰富情感内容。TextrolSpeech的应用领域主要集中在提高TTS系统的风格控制能力和情感表达,特别是在情感效果的领域。
TextrolSpeech is a large-scale open-source speech emotion dataset developed by Zhejiang University, which contains 330 hours of speech data and 236,220 pairs of natural text-style prompts. Generated via a multi-stage prompt programming approach, this dataset covers five stylistic factors including gender, pitch, speech rate, volume, and emotion, aiming to advance the development of text-controllable TTS systems. During its construction, LibriTTS and VCTK datasets were integrated, and multiple additional emotion datasets were collected to enrich the emotional content. The main application fields of TextrolSpeech focus on improving the style control and emotional expression capabilities of TTS systems, especially in the domain of emotional speech effects.




