EmoSpeech
收藏资源简介:
EmoSpeech数据集是由香港科技大学和香港浸会大学联合创建的情感丰富且上下文详细的语音标注语料库。该数据集包含约16小时的音频,主要从电影和电视剧中提取,涵盖了多种情感表达和场景。每个样本都通过自然语言句子进行详细描述,而非传统的固定情感标签,为情感控制的文本到语音(TTS)系统提供了更准确的数据。数据集的创建过程包括目标语音提取、情感识别和数据增强,利用生成模型和大型语言模型(LLM)进行自动标注和数据扩充,减少了手动标注的成本。该数据集的应用领域主要集中在情感控制的TTS系统开发,旨在解决现有情感语音数据库标注简单、情感表达不足的问题。
The EmoSpeech dataset is a richly emotional and contextually detailed speech annotated corpus jointly created by Hong Kong University of Science and Technology and Hong Kong Baptist University. It contains approximately 16 hours of audio, mainly extracted from films and TV dramas, covering diverse emotional expressions and scenarios. Each sample is elaborately described using natural language sentences rather than traditional fixed emotion labels, providing more accurate data for emotion-controlled text-to-speech (TTS) systems. The dataset creation process includes target speech extraction, emotion recognition and data augmentation, where generative models and large language models (LLMs) are employed for automatic annotation and data expansion, thereby reducing the cost of manual annotation. The main application fields of this dataset focus on the development of emotion-controlled TTS systems, aiming to solve the problems of simplistic annotation and insufficient emotional expression in existing emotional speech databases.




