character-voices
收藏资源简介:
Character Voices — DramaBox-annotated Echo-TTS 是一个高质量、情感条件化的角色语音数据集,专为文本到语音(TTS)任务设计。该数据集包含由 Echo-TTS 生成的语音,并经过精心处理:使用 Parakeet-TDT-0.6B-v3 进行转录,经过静音感知修剪,过滤至词错误率(WER)为 0.00,并通过 RE-USE 和 LavaSR 技术增强至 48 kHz 采样率。每个语音样本都使用 Gemini 模型标注了详细的 DramaBox 表演提示,包括角色原型描述、可感知的情感和表达方式,以及精确的台词(用双引号标出)。数据集共包含 14 个风格各异的虚构角色(如精灵、怪物、僵尸、机器人、龙等),总计 2184 个语音样本。每个角色对应一个独立的文件夹和一个 WebDataset tar 归档文件。每个样本由一对文件组成:一个 160 kbps 的单声道 MP3 音频文件和一个包含丰富元数据的 JSON 文件。JSON 元数据字段包括:转录文本(transcription)、自动语音识别信息(asr)、DramaBox 表演提示(dramabox_prompt)、感知到的情感(perceived_emotions)、角色原型描述符(archetype_descriptor)以及情感标签(emotion)。该数据集适用于需要高保真、情感丰富且具有特定角色特征的语音合成模型训练、评估和研究,特别是在角色扮演、游戏、有声内容创作和情感 TTS 等领域。
Character Voices — DramaBox-annotated Echo-TTS is a high-quality, emotion-conditioned character voice dataset designed for text-to-speech (TTS) tasks. The dataset contains speech generated by Echo-TTS, meticulously processed: transcribed using Parakeet-TDT-0.6B-v3, silence-aware trimmed, filtered to a word error rate (WER) of 0.00, and enhanced to a 48 kHz sampling rate via RE-USE and LavaSR technologies. Each speech sample is annotated with detailed DramaBox performance prompts using the Gemini model, including character archetype descriptions, perceived emotions and expressions, and precise lines (enclosed in double quotes). The dataset includes 14 diverse fictional characters (such as elves, monsters, zombies, robots, dragons, etc.), totaling 2184 speech samples. Each character corresponds to a separate folder and a WebDataset tar archive file. Each sample consists of a pair of files: a 160 kbps mono MP3 audio file and a JSON file containing rich metadata. The JSON metadata fields include: transcription text (transcription), automatic speech recognition information (asr), DramaBox performance prompts (dramabox_prompt), perceived emotions (perceived_emotions), character archetype descriptors (archetype_descriptor), and emotion labels (emotion). This dataset is suitable for training, evaluating, and researching high-fidelity, emotion-rich speech synthesis models with specific character traits, particularly in areas such as role-playing, gaming, audio content creation, and emotional TTS.
数据集概述
- 数据集名称: Character Voices — DramaBox-annotated Echo-TTS (WER=0)
- 许可证: cc-by-nc-sa-4.0
- 任务类别: 文本到语音 (Text-to-Speech)
- 语言: 英语 (en)
数据生成与处理流程
- 使用 Echo-TTS 生成情感条件化的角色语音。
- 通过 Parakeet-TDT-0.6B-v3 进行转录。
- 进行静音感知修剪。
- 过滤至词错误率(WER)为 0.00。
- 使用 RE-USE 和 LavaSR 进行 48 kHz 音频增强。
- 通过 Gemini 模型进行标注,生成 DramaBox 表演提示(包含角色原型、可听情感、表达方式及精确台词)。
数据格式与结构
- 每个角色对应一个单独的文件夹和一个 WebDataset tar 包 (
wds/<character>.tar)。 - 每个样本包含:
<key>.mp3:160 kbps 单声道音频文件。<key>.json:包含transcription、asr、dramabox_prompt、perceived_emotions、archetype_descriptor、emotion字段的 JSON 文件。
角色与样本数量
总共包含 14 个角色,共计 2184 个样本。
| 角色名称 | 样本数量 |
|---|---|
| Fairy-2 | 153 |
| Fluffy_Cookie_Monster | 125 |
| Goblin - en | 162 |
| Old noble Dragon | 181 |
| Zombie-Chris2 | 136 |
| Zombie-Chris3 | 147 |
| charming flity woman | 200 |
| clichee-goblin | 166 |
| cuddle-gnome | 176 |
| monsterous-orc | 169 |
| cute cartoon animal | 172 |
| robot | 101 |
| cartoon gnome | 151 |
| zombie-ref | 145 |




