SpeechCraft
收藏资源简介:
SpeechCraft是由清华大学创建的一个细粒度的双语表达性语音数据集,旨在促进语音与自然语言的多模态学习。该数据集包含约2,000小时的音频数据和超过两百万条语音片段,通过自动语音标注系统生成详细的自然语言描述。数据集的创建过程结合了专家分类器和高级标题模型,以及经过微调的LLaMA模型,以捕捉和描述语音的细微特性。SpeechCraft主要应用于可控语音生成和自动化语音标题生成领域,旨在提高语音合成和语音风格理解的性能。
SpeechCraft is a fine-grained bilingual expressive speech dataset developed by Tsinghua University, aiming to advance multimodal learning between speech and natural language. This dataset contains approximately 2,000 hours of audio data and over two million speech segments, with detailed natural language descriptions generated via an automatic speech annotation system. The dataset creation process integrates expert classifiers, advanced speech captioning models, and fine-tuned LLaMA models to capture and describe the subtle characteristics of speech. SpeechCraft is primarily applied in the domains of controllable speech generation and automated speech captioning, with the goal of improving the performance of speech synthesis and speech style understanding.

- 1SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description清华大学 · 2024年



