moss-local-voice-acting-v4-emorant-64x200-part2
收藏资源简介:
MOSS-Local Voice Acting v4 emorant(第二部分/共四部分)是一个专注于生成高情感表现力语音的文本到语音(TTS)合成数据集。该数据集是更大规模语料库的一部分,旨在通过精心设计的提示语,推动语音表达超越中性朗读,达到极强的情感强度。数据内容包括合成语音音频及其丰富的元数据,本部分涵盖10种情感类别:满足、失望、厌恶、痛苦、怀疑、兴高采烈、尴尬、情感麻木、疲劳/筋疲力尽、恐惧。每种情感对应200组,每组包含64个不同的语音生成片段(takes),采用48 kHz单声道FLAC格式。整个四部分语料库的总体目标是生成512,000个音频片段。语音由合并的4.55B参数本地Transformer架构的MOSS-TTS语音表演模型生成,温度为1.0,直接使用原始模型输出,关键特征为“无参考音频”,即模型完全基于文本角色描述创造声音,而非模仿现有音频。数据生成过程经过特殊设计以最大化情感表达:每种情感映射到特定“表达类别”(如失控的咆哮、悲痛、喜悦、温柔等),并使用类别相关的声音机制描述词(如撕裂声、哽咽、抽泣、欢呼、耳语等)。提示语以具象化方式命名感受,构建情感升级叙事弧,在脚本中嵌入括号内非言语声音爆发提示,并通过结尾保护性指令确保语音可懂度。所有脚本均为全新的英文提示。数据集提供完整元数据,每个情感压缩包包含scores.parquet评分文件和cells.jsonl元数据文件。评分文件为每个take提供详细数据,包括组ID、目标情感、共情洞察力头部、提示索引、种子值、声音爆发提示;基于ASR的词错误率(WER)及其倒数、识别文本;VoiceCLAP商业模型的情感混合度、真诚度评分,以及音频与风格指令的文本余弦相似度;共情洞察力语音增强模型的40种情感维度分数、唤醒度、效价和真实性评分;音频响度、持续时间、生成是否正常结束标志;以及本组所提示目标情感的目标分数。数据集未经任何过滤,包含所有生成片段,允许用户根据自身需求(如结合情感评分、真诚度和可懂度)实施自定义的“最佳N选择”策略。适用于高表现力TTS模型的研究、训练与评估,情感语音生成分析,以及语音合成输出质量的选择与优化等任务。数据集采用Apache-2.0许可证,所有音频均为完全合成生成。
MOSS-Local Voice Acting v4 emorant (Part 2 of 4) is a text-to-speech (TTS) synthesis dataset focused on generating highly emotionally expressive speech. It is part of a larger corpus designed to push speech expression beyond neutral reading to achieve intense emotional intensity through carefully crafted prompts. The data consists of synthetic speech audio and rich metadata. This part includes 10 specific emotional categories: satisfaction, disappointment, disgust, pain, doubt, elation, embarrassment, emotional numbness, fatigue/exhaustion, fear. Each emotion corresponds to 200 groups, with each group containing 64 different speech generation takes, in 48 kHz mono FLAC format. The overall goal of the entire four-part corpus is to generate 512,000 audio segments. The speech is generated by a merged 4.55B parameter local Transformer architecture MOSS-TTS voice acting model at a temperature of 1.0, using raw model outputs directly. A key feature is no reference audio, meaning the model creates voices entirely based on textual character descriptions of a generic human speaker, rather than imitating existing audio. The data generation process is specially designed to maximize emotional expression: each emotion is mapped to a specific expression category (e.g., uncontrolled roar, grief, joy, tenderness) and uses category-related vocal mechanism descriptors (e.g., tearing, choking, sobbing, cheering, whispering). Prompts name feelings in a concrete manner, build emotional escalation narrative arcs, embed non-verbal vocal burst cues in parentheses within scripts, and ensure speech intelligibility through end-protective instructions. All scripts are new English prompts. The dataset provides complete metadata. Each emotion archive includes a scores.parquet scoring file and a cells.jsonl metadata file in addition to audio files. The scoring file provides a detailed row for each take, including: group ID, target emotion, used empathy insight head, prompt index, seed value, vocal burst cue; ASR-based word error rate (WER) and its reciprocal, recognized text; VoiceCLAP commercial models emotional blend, sincerity scores, and text cosine similarity between audio and style instructions; empathy insight voice enhancement models 40 emotional dimension scores, arousal, valence, and authenticity scores; audio loudness, duration, flag for normal generation completion; and target score for the prompted target emotion of the group. The dataset is unfiltered and includes all generated takes, allowing users to implement custom best N selection strategies based on their needs (e.g., combining emotion scores, sincerity, and intelligibility). It is suitable for research, training, and evaluation of highly expressive TTS models, emotional speech generation analysis, and selection and optimization of speech synthesis output quality. The dataset is licensed under Apache-2.0, and all audio is fully synthetic.
数据集概述:MOSS-Local Voice Acting v4 "emorant" (part 2/4)
基本信息
- 数据集名称:MOSS-Local Voice Acting v4 "emorant" (part 2/4)
- 许可证:Apache-2.0
- 任务类别:文本转语音(Text-to-Speech)
- 语言:英语
- 数据集大小:包含40种情绪 × 200组 × 64条录音(4部分总计目标512,000个音频片段)
- 生成模型:基于合并的4.55B参数本地变换器MOSS-TTS语音表演模型,原生48kHz采样率,温度1.0,未经滤波的原始模型输出,无参考音频(模型根据普通人类说话者的人物描述自主创造声音)
数据集目的与设计
v4版本的核心目的是情感强度:指令来自于一项关于情感原型的提示研究,旨在将语音输出推向远超出中性朗读式语音的极端表达:
- 每种情绪映射到一个表演类别(失控咆哮、痛苦、悲伤、喜悦、温柔、惊奇,或特殊类别如醉酒/戏弄/狂喜),并使用类别特定的语音机制语言(撕裂、破音、啜泣、高呼、低语等)
- 情感以身体感受的方式命名,指令构建逐步升级的叙事弧(“每一行都更加恶毒”)
- 脚本中嵌入语音爆发提示
- 末尾使用保证语(“……但每个词依然清晰可辨”)维持可懂度
本部分包含的情绪(10种)
- Contentment(满足)
- Disappointment(失望)
- Disgust(厌恶)
- Distress(痛苦)
- Doubt(怀疑)
- Elation(兴高采烈)
- Embarrassment(尴尬)
- Emotional_Numbness(情感麻木)
- Fatigue_Exhaustion(疲劳/筋疲力尽)
- Fear(恐惧)
数据内容与格式
每个 data/<Emotion>.tar 压缩包包含:
- 64条录音组:
<Emotion>_<group>_v<take>.flac(48kHz单声道FLAC格式) scores.parquet:评分文件,每行为一条录音cells.jsonl:包含每组的确切文本、指令和生成元数据
scores.parquet 字段说明
| 字段 | 描述 |
|---|---|
gid, emotion, ei_head, prompt_idx, seed |
组ID、目标情绪、使用的共情洞察头部、提示索引、录音索引 |
vocal_burst |
指令中包含的非语言爆发提示 |
wer, inv_wer, hyp |
与脚本的字错误率(使用Parakeet-TDT-0.6b-v3 ASR)、1/(1+WER)和ASR转录文本 |
blend, genu, prompt_sim |
VoiceCLAP-commercial情感混合评分、真实性评分、音频-文本与风格指令的余弦相似度 |
ei_*(42列) |
共情洞察-语音-Plus评分:40个情感维度 + ei_Arousal、ei_Valence(以及 ei_Authenticity) |
rms_db, peak_db, dur, finished |
响度、峰值、时长(秒),以及生成是否以EOS(而非令牌上限)结束 |
target |
该组被提示的情感EI-Plus评分 |
无任何过滤:所有录音均包含在内,用户可自行应用自己的最优选择策略。
相关资源
- 生成模型:
laion/moss-tts-local-transformer-4.55b-voice-acting - 推理代码与指南:LAION-AI/laion-moss-local-1.5-voice-acting-4.55b
- 同类发布:v2无参考语料库
laion/moss-local-voice-acting-64x100-part1..4,完整DramaBox再解释版本laion/moss-local-dramabox-full-reinterpretations-64
许可证
Apache-2.0。所有音频均为完全合成生成(无参考集不包含任何真实说话者的声音)。




