gemini-2.5-pro-tts-voice-profiles-dramabox
收藏资源简介:
该数据集是原始laion/gemini-2.5-pro-tts-voice-profiles数据集的增强版本,专门用于文本到语音(TTS)任务,特别关注语音表演、情感表达和语音配置文件生成。核心改进在于为每个样本新增了两个由google/gemma-3-12b-it模型生成的DramaBox风格表演提示词,这些提示词基于样本原有的标注(如BUD-E Whisper转录、Empathic-Insight情感分数、Gemini词级转录/描述和发声行为描述)生成,避免了重新运行自动语音识别(ASR)或Whisper模型的需求。每个样本的JSON文件在保留原始所有字段的基础上,新增了以下字段:1) `dramabox_prompt_full`:完整的DramaBox表演提示词,包含说话者的人口统计信息(如年龄、性别、音色)、情感舞台指示、对话,并将发声行为(如呼吸声、叹息)融入舞台指示中;2) `dramabox_prompt_focused`:聚焦的DramaBox提示词,不包含人口统计信息,仅关注音质、情感表达和对话,适用于需要参考音频或无偏TTS的场景;3) `dramabox_model`:生成提示词的模型ID;4) `dramabox_meta.full_starter` / `focused_starter`:用于增加提示词开头多样性的随机起始词或短语。提示词遵循固定格式:首先是一个以...delivering this high-quality studio voice recording with no background noise.结尾的说话者/声音描述段落,随后是交替出现的`(舞台指示)`和`"对话"`。示例展示了多样化的说话者角色、复杂的情感状态(如优越感/悲伤、愠怒/嫉妒、羞怯/厌恶、愉悦/满足、勇气/失望等)以及丰富的非语言发声行为(如咳嗽、啜泣等)。该数据集旨在为高质量、富有表现力的语音合成提供丰富的、带有详细表演指导的文本输入,适用于需要控制语音情感、风格和表演细节的TTS研究和应用。
This dataset is an enhanced version of the original laion/gemini-2.5-pro-tts-voice-profiles dataset, specifically designed for text-to-speech (TTS) tasks, with a focus on voice performance, emotional expression, and voice profile generation. The core improvement involves adding two DramaBox-style performance prompts generated by the google/gemma-3-12b-it model for each sample. These prompts are based on the samples original annotations (such as BUD-E Whisper transcriptions, Empathic-Insight emotion scores, Gemini word-level transcriptions/descriptions, and vocal behavior descriptions), eliminating the need to rerun automatic speech recognition (ASR) or Whisper models. Each samples JSON file retains all original fields and includes new fields: 1) `dramabox_prompt_full`: a complete DramaBox performance prompt containing speaker demographic information (e.g., age, gender, timbre), emotional stage directions, dialogue, and integrating vocal bursts (e.g., breathing, sighs) into the stage directions; 2) `dramabox_prompt_focused`: a focused DramaBox prompt without demographic information, emphasizing only voice quality, emotional expression, and dialogue, suitable for scenarios requiring reference audio or unbiased TTS; 3) `dramabox_model`: the model ID used to generate the prompts; 4) `dramabox_meta.full_starter` / `focused_starter`: random starting words or phrases to increase prompt diversity. Prompts follow a fixed format: starting with a speaker/voice description paragraph ending with ...delivering this high-quality studio voice recording with no background noise., followed by alternating `(stage directions)` and `"dialogue"`. Examples showcase diverse speaker roles, complex emotional states (e.g., superiority/sadness, resentment/jealousy, shyness/disgust, joy/satisfaction, courage/disappointment), and rich non-verbal vocal behaviors (e.g., coughing, sobbing). The dataset aims to provide rich, detailed performance-guided text input for high-quality, expressive speech synthesis, applicable to TTS research and applications that require control over voice emotion, style, and performance details.
数据集名称
Gemini 2.5 Pro TTS Voice Profiles — with DramaBox prompts
许可证
CC-BY-4.0
任务类别
文本转语音(text-to-speech)
标签
voice-acting, emotion, gemini-tts, dramabox, voice-profiles
数据集描述
该数据集是 laion/gemini-2.5-pro-tts-voice-profiles 的增强版本。原始数据集中的每个样本都通过 google/gemma-3-12b-it 模型,利用已有的标注(BUD-E Whisper 字幕、Empathic-Insight 分数、Gemini 词级转录/字幕以及发声爆发描述)生成了两个 DramaBox 性能提示,无需重新运行 ASR 或 Whisper 专家模型。
新增字段
每个样本的 .json 文件新增了以下字段:
| 字段 | 描述 |
|---|---|
dramabox_prompt_full |
包含说话者人口统计学信息(年龄、性别、音色)+ 情感舞台指令 + 对话的 DramaBox 性能提示。发声爆发被编织进舞台指令中。 |
dramabox_prompt_focused |
不包含人口统计学信息的 DramaBox 提示——仅保留声音质量 + 情感表达 + 对话(适用于参考音频/无偏 TTS)。 |
dramabox_model |
生成该提示的模型 ID。 |
dramabox_meta.full_starter / focused_starter |
用于开场白的随机单词或短语(确保开场多样性,避免所有提示以相同方式开头)。 |
提示格式
提示以一段说话者/声音描述段落结尾,内容为“delivering this high-quality studio voice recording with no background noise.”,之后交替出现 (stage directions) 和 "dialogue"。
示例提示
数据集 README 中提供了多个示例,包括带有说话者人口统计学信息的“FULL”提示和不带人口统计学信息的“FOCUSED”提示,涵盖了不同的情感(如 superiority/mournfulness、sullenness/jealousy & envy、pleasure/fulfillment、courage/disappointment 等)和发声爆发(如 breath_catch、teeth_click、pain_gasp、pleading_whine 等)。




