LibriTTS-P
收藏资源简介:
LibriTTS-P是由日本LY Corp.创建的一个基于LibriTTS-R的新型语音数据集,专注于提供详细的语音风格和说话者身份提示。该数据集包含373,868条记录,通过混合方法构建提示注释,包括手动和合成注释,以捕捉人类对说话者特征的感知和语音风格。数据集的创建过程涉及对基本频率、每秒音节数和响度的统计分析,以及使用大型语言模型进行数据增强。LibriTTS-P主要应用于基于提示的可控文本到语音转换(TTS)和风格标题生成,旨在提高TTS模型的自然度和风格描述的准确性。
LibriTTS-P is a novel speech dataset developed by LY Corp. of Japan, which is built upon LibriTTS-R and focuses on providing detailed speech style and speaker identity prompts. This dataset comprises 373,868 records, and its prompt annotations are constructed through a hybrid methodology integrating manual and synthetic annotations to capture human perceptions of speaker traits and speech styles. The development process of this dataset includes statistical analyses of fundamental frequency, syllables per second, and loudness, as well as data augmentation employing large language models (LLMs). LibriTTS-P is mainly utilized for prompt-based controllable text-to-speech (TTS) and style caption generation, with the objective of improving the naturalness of TTS models and the accuracy of style descriptions.




