laion/character-voices
收藏资源简介:
该数据集是一个角色语音数据集,包含14个不同角色(如精灵、怪物、机器人等)的2184个情感条件语音样本。语音通过Echo-TTS生成,使用Parakeet-TDT-0.6B-v3进行转录,经过静音感知修剪,过滤至词错误率(WER)为0.00,并采用RE-USE和LavaSR 48 kHz技术增强。每个样本都通过Gemini标注为DramaBox表演提示,包括角色原型、可听情感、表达方式和具体台词(用双引号标注)。数据以每个角色一个文件夹和一个WebDataset tar文件的形式组织,每个样本包含一个MP3音频文件(160 kbps单声道)和一个JSON文件,其中JSON文件包含转录文本、自动语音识别结果、DramaBox提示、感知情感、原型描述器和情感标签。
This dataset is a character voices dataset containing 2184 emotion-conditioned speech samples across 14 different characters (e.g., fairy, monster, robot). The speech is generated using Echo-TTS, transcribed with Parakeet-TDT-0.6B-v3, silence-aware trimmed, filtered to a word error rate (WER) of 0.00, and enhanced using RE-USE and LavaSR 48 kHz. Each sample is annotated with Gemini into DramaBox performance prompts, including archetype, audible emotions, delivery, and the exact line in double quotes. The data is organized with one folder and one WebDataset tar file per character, where each sample includes an MP3 audio file (160 kbps mono) and a JSON file containing transcription, ASR, DramaBox prompt, perceived emotions, archetype descriptor, and emotion.




