dramabox-cutscene-prompts
收藏资源简介:
DramaBox Cut-Scene Voice-Acting Prompts是一个持续生成的、角色一致的文本数据集,专门用于训练和评估表达性文本到语音(TTS)及语音表演模型。数据集包含909,000个提示,每个提示描述一个说话者在两个情感对比强烈的场景之间通过“CUT TO:”转换,采用DramaBox舞台指示格式(口语内容用引号标注,表演说明用括号标注)。数据仅包含文本,无音频。提示从Voice-Acting-Pipeline分类法中采样,涵盖5种CUT-TO路径:CCA(基于VoiceNet声学属性)、CC2-C(基于类型原型和节奏/唤醒度)、ACCC(基于表演挑战简报)、SIT(基于情境)和Extreme Physical(从平静感官性爆发为极端身体感觉)。每个系统提示包含情感强调指令(要求情感原始、对比强烈)和一个简洁的声爆发菜单(SFW分类法,180种爆发),使模型能自然插入非语音爆发作为导演说明。约50%的提示随机注入一个声爆发(记录在injected_burst字段)。对话长度默认约50个口语词(每场景约25词),每个生成独立采样路径、语言、分类属性和词种子,确保输入状态唯一。数据分布包括语言(英语454,005条,德语454,995条)、路径、感知性别(女性530,011条,男性291,598条,其他87,391条)、年龄组(如年轻成人288,455条)以及前15种情感(如悲伤、自豪、解脱)和注入声爆发(如哀悼挽歌、鼻哼声)。数据模式包含多个字段,关键字段有dramabox_prompt(生成的文本)、路径、语言、采样属性(性别、年龄组、唤醒度、情感等)、词种子、注入声爆发、生成设置(模型、温度、最大令牌数)和令牌计数。数据集使用CC-BY-4.0许可证,由本地Gemma模型生成,旨在支持表达性语音表演和TTS研究。
DramaBox Cut-Scene Voice-Acting Prompts is a continuously generated, character-consistent text dataset specifically designed for training and evaluating expressive text-to-speech (TTS) and voice performance models. The dataset contains 909,000 prompts, each describing a speaker transitioning between two emotionally contrasting scenes via a "CUT TO:" transition, following the DramaBox stage direction format where spoken content is enclosed in quotation marks and performance instructions are enclosed in parentheses. All data is text-only, with no audio content included. The prompts are sampled from the Voice-Acting-Pipeline taxonomy, covering 5 CUT-TO path categories: CCA (based on VoiceNet acoustic attributes), CC2-C (based on genre prototypes and arousal/valence), ACCC (based on performance challenge briefings), SIT (based on context), and Extreme Physical (transitioning from calm sensory experiences to extreme physical sensations). Each system prompt incorporates emotion-focused instruction requirements (demanding raw, strongly contrasting emotions) and a concise vocal burst menu (SFW taxonomy, 180 distinct bursts), enabling models to naturally insert non-verbal vocal bursts as director’s notes. Approximately 50% of prompts randomly inject a vocal burst, which is recorded in the "injected_burst" field. The default dialogue length is approximately 50 spoken words (around 25 words per scene), with each generation independently sampling paths, languages, classification attributes, and word seeds to ensure unique input states. The dataset’s distribution covers multiple dimensions: languages (454,005 entries in English, 454,995 in German), CUT-TO paths, perceived gender (530,011 female, 291,598 male, 87,391 other), age groups (e.g., 288,455 entries for young adults), the top 15 emotions (e.g., sadness, pride, relief), and injected vocal bursts (e.g., lament, nasal snort). The dataset schema includes multiple fields, with key fields including "dramabox_prompt" (the generated text), path, language, sampling attributes (gender, age group, arousal, emotion, etc.), word seed, injected vocal burst, generation settings (model, temperature, maximum token count), and token count. The dataset is licensed under CC-BY-4.0, generated using a local Gemma model, and aims to support research on expressive voice performance and TTS.
数据集概述
数据集名称: DramaBox Cut-Scene Voice-Acting Prompts
数据集地址: https://huggingface.co/datasets/laion/dramabox-cutscene-prompts
许可证: CC-BY-4.0
描述
该数据集是一个纯文本数据集,不包含音频。它包含连续生成的、角色一致的两场景“CUT TO:” 语音表演提示,用于训练和评估表现力丰富的文本转语音(TTS)或配音模型。每个提示描述一个说话者在两个情绪对比强烈的时刻,并由 CUT TO: 分隔,采用 DramaBox 舞台指示格式(对话用引号,表演说明用括号)。
数据规模与语言
- 提示总数: 1,514,000 条
- 语言: 英语和德语
- 语言分布: 德语 757,443 条,英语 756,557 条
生成方式
提示通过 Voice-Acting-Pipeline 分类体系采样生成,沿以下 5 种 CUT-TO 路径(pathway)产生具有强烈情绪对比的两场景提示:
- CCA (VoiceNet) — 基于采样的 VoiceNet 声学属性构建说话者
- CC2-C (Archetype) — 体裁原型 + 节奏 + 唤醒度
- ACCC (Acting Challenge) — 表演挑战简报
- SIT (Situations) — 说话者身处影响声音的情境中
- Extreme Physical — 从平静的感官体验爆发为极端身体感受
生成的系统提示包含情绪强调指令(使其更原始、将对比推至极限)和紧凑的声爆发声菜单(SFW 分类,180 种爆发声)。约 50% 的行会随机注入一个必需的声音爆发元素,并记录在 injected_burst 字段中。对话长度遵循默认设置(约 50 个口语词,每场景约 25 个词)。每次生成独立采样路径、语言、分类属性以及通过 os.urandom 强随机性生成的种子词。
- 声爆发声分类体系: LAION vocal_bursts_taxonomy_sfw.json(LAION-AI/voice-taxonomies vocalburst 的 SFW 子集,版本
sfw-v1) - 爆发声注入: 757,676 条包含必需爆发声,756,324 条为自由选择
数据分布
按路径/条件
| 路径 | 数量 |
|---|---|
| Extreme Physical | 303,321 |
| CCA (VoiceNet) | 303,160 |
| ACCC (Acting Challenge) | 303,028 |
| SIT (Situations) | 302,267 |
| CC2-C (Archetype) | 302,224 |
按感知性别
| 性别 | 数量 |
|---|---|
| 女性 | 882,371 |
| 男性 | 486,110 |
| 其他 | 145,519 |
按年龄组
| 年龄组 | 数量 |
|---|---|
| 年轻成人 | 480,918 |
| 中年 | 317,190 |
| 老年 | 283,787 |
| 未指定 | 232,185 |
| 青少年 | 197,313 |
| 成人 | 2,607 |
主要情绪(前15名)
| 情绪 | 数量 |
|---|---|
| 悲伤 (Sadness) | 47,204 |
| 自豪 (Pride) | 47,198 |
| 如释重负 (Relief) | 47,192 |
| 满足 (Contentment) | 47,159 |
| 困惑 (Confusion) | 46,367 |
| 希望/乐观 (Hope/Optimism) | 46,301 |
| 尴尬 (Embarrassment) | 46,287 |
| 厌恶 (Disgust) | 46,249 |
| 醉酒/意识状态改变 (Intoxication/Altered States) | 46,212 |
| 疲劳/精疲力竭 (Fatigue/Exhaustion) | 46,164 |
| 羞耻 (Shame) | 46,152 |
| 胜利 (Triumph) | 46,057 |
| 失望 (Disappointment) | 45,887 |
| 苦涩 (Bitterness) | 45,841 |
| 愉快 (Amusement) | 45,787 |
主要注入的声爆发声(前15名)
| 爆发声 | 数量 |
|---|---|
| Breathy Oh no | 4,369 |
| Relaxing Exhale | 4,360 |
| Nose-Wrinkling Snort | 4,349 |
| Aha! Realization Burst | 4,337 |
| Savoring Mmm | 4,335 |
| Chortle | 4,326 |
| Sputtering/Spitting | 4,313 |
| Blech Vocalization | 4,308 |
| Busking / Street Singing | 4,304 |
| Humming While Working / Cooking | 4,300 |
| Mourning Dirge / Funeral Hymn | 4,298 |
| Breathless Laugh | 4,296 |
| Anticipatory Lip Smack | 4,295 |
| Conversational Uh-uh (No) | 4,295 |
| Wolf-Whistle | 4,294 |
数据模式 (Schema)
每一行(一个生成的提示)包含以下字段:
uid, global_index, gen_seed, worker_gpu, pathway, pathway_label, lang, dramabox_prompt, gender, age_group, sampled_gender, sampled_age, arousal, emotions, voicenet_attributes, flow_style, emotion_alignment, direction_style, archetype, genre, tempo, situation, situation_dim, challenge_title, challenge_id, category, subcategory, word_seeds, n_word_seeds, injected_burst, burst_taxonomy_source, burst_taxonomy_version, model, precision, temperature, top_p, max_tokens, input_tokens, output_tokens, gen_timestamp
关键字段:
dramabox_prompt— 生成的文本内容pathway/lang— 生成路径和语言gender、age_group、arousal、emotions、voicenet_attributes、archetype/genre、situation、challenge_title、category— 采样属性word_seeds、injected_burst— 种子词和注入的爆发声model、temperature、max_tokens— 生成设置input_tokens/output_tokens— 输入和输出 token 数uid和global_index— 唯一标识符
生成模型参数
- 模型:
unsloth/gemma-3n-E4B-it(bf16) - 温度: 0.9
- 最大 token 数: 1024
- 平均 token 数: 输入 3251 / 输出 368
数据分片
数据以 Parquet 文件形式存储,每文件 1,000 行,位于 data/ 目录下。
用途与许可
该数据集旨在用于表现力丰富的配音 / TTS 的研究。它由本地 Gemma 模型生成的合成数据构成,不包含任何音频。采用 CC-BY-4.0 许可。




