dramabox-burst-prompts
收藏资源简介:
DramaBox 场景提示数据集专为文本到语音合成中的非语言发声(vocal bursts)设计,包含 24,500 个脚本风格场景提示,均匀分布在 49 个发声类别中(每个类别 500 个)。每个提示都通过精心设计的结构确保爆发在正常对话中自然出现,而非孤立音效:提示以平静开场(包含明确的正常语速指令和一句日常对话),随后插入一个非语言发声,大部分场景(62%)在爆发后恢复平静。该数据集旨在修复先前场景中因缺乏平静引入而导致的动态对比不足问题,使模型能够区分“此处有爆发”与“整个话语都很响亮”。数据由 google/gemma-4-31B-it 生成,并通过结构化组合(地点、话题、触发因素、说话者人口统计等)保证多样性。所有提示均经过质量门控(0.9% 拒绝率),确保爆发拼写正确、平静开场符合要求。数据集包含 16 个字段,包括发声类别、强度(loud/mid/quiet)、爆发拼写、说话者性别(男女各半)、年龄组(青年/中年/老年)、地点、触发因素、爆发位置等。提示文本采用 CC BY 4.0 许可,可自由使用;但若使用 DramaBox 模型生成音频,则需遵守 LTX-2 社区许可(可能限制商业用途)。该数据集适用于训练 TTS 模型生成自然融入语音的非语言发声,或用于训练爆发检测器与适配器。
The DramaBox Scene Prompt Dataset is designed for non-verbal vocal bursts in text-to-speech synthesis, containing 24,500 script-style scene prompts uniformly distributed across 49 vocal burst categories (500 per category). Each prompt is carefully structured to ensure the burst appears naturally within normal conversation rather than as an isolated sound effect: the prompt begins calmly (with an explicit instruction for normal speech rate and a daily conversation sentence), then inserts a non-verbal vocal burst, with most scenes (62%) returning to calm after the burst. The dataset aims to address the lack of dynamic contrast caused by the absence of calm introduction in previous scenes, enabling the model to distinguish between there is a burst here and the entire utterance is loud. Data is generated by google/gemma-4-31B-it and diversity is ensured through structured combinations (location, topic, trigger, speaker demographics, etc.). All prompts undergo quality gating (0.9% rejection rate) to ensure correct burst spelling and compliant calm opening. The dataset includes 16 fields, including vocal burst category, intensity (loud/mid/quiet), burst spelling, speaker gender (balanced male/female), age group (young/middle-aged/elderly), location, trigger, burst position, etc. The prompt text is licensed under CC BY 4.0 for free use; however, if using the DramaBox model to generate audio, the LTX-2 Community License must be followed (which may restrict commercial use). The dataset is suitable for training TTS models to generate non-verbal vocal bursts naturally integrated into speech, or for training burst detectors and adapters.




