moss-local-voice-acting-v4-emorant-64x200-part4
收藏资源简介:
MOSS-Local Voice Acting v4 "emorant"(第四部分/共四部分)是一个专注于高情感强度的语音合成数据集,属于四部分语料库的最终部分,旨在生成超越中性朗读的、充满情感表现力的语音样本。数据使用4.55B参数本地Transformer MOSS-TTS语音表演模型生成,采样率为48kHz,不依赖参考音频,仅基于人物描述创造声音。v4版本通过重写提示指令,将情感映射到特定表达类别(如咆哮、痛苦、喜悦等),并嵌入声音爆发提示,以增强情感强度。数据集包含10种情感:愉悦_狂喜、自豪、解脱、悲伤、性欲_渴望、羞耻、酸楚、戏弄、感谢_感激、胜利。每种情感对应一个.tar文件,内含64个48kHz单声道FLAC格式录音、scores.parquet文件(提供组ID、情感、分数等元数据)和cells.jsonl文件(包含文本、指令和生成元数据)。数据集未经过滤,适用于文本到语音合成、情感语音生成、语音表现力分析和语音情感计算等研究任务。
MOSS-Local Voice Acting v4 "emorant" (Part 4 of 4) is a speech synthesis dataset focused on high emotional intensity, serving as the final part of a four-part corpus. It aims to generate emotionally expressive speech samples beyond neutral reading. The data is generated using a merged 4.55B parameter local Transformer MOSS-TTS voice acting model at a 48kHz sample rate, relying solely on character descriptions of ordinary human speakers without any reference audio. The core goal of v4 is emotional intensity: prompts have been rewritten based on research into emotional archetypes to push vocal expression to extremes. Strategies include mapping each emotion to specific expression categories (e.g., uncontrolled roaring, pain, sadness, joy, tenderness, surprise, and special categories for intoxication/teasing/ecstasy) with category-specific vocal mechanism descriptions (e.g., tearing, cracking, sobbing, cheering, whispering); emotion names are intuitive and intense, with instructions building escalating emotional arcs; scripts embed vocal burst cues in brackets; and safeguards are included to maintain intelligibility. The dataset includes 10 emotions: joy_ecstasy, pride, relief, sadness, lust_longing, shame, bitterness, teasing, gratitude_thanks, triumph. Data organization: each emotion corresponds to a .tar file containing 64 recordings (48kHz mono FLAC), a scores.parquet file, and a cells.jsonl file (with exact text, instructions, and generation metadata for each group). The scores.parquet file provides detailed per-recording data, including group ID, target emotion, Empathic-Insight head used, prompt index, seed, vocal burst cues, word error rate with ASR transcript, VoiceCLAP commercial emotion blend score, authenticity score, audio-text style instruction cosine similarity, 42 Empathic-Insight-Voice-Plus emotional dimension scores (including arousal and valence), loudness, duration, generation completion flag, and EI-Plus score for the groups prompt target emotion. The dataset is unfiltered, containing all generated recordings, allowing users to perform best-N selection based on custom criteria (e.g., combining normalized emotion blend scores, authenticity scores, and inverse word error rate). It is suitable for text-to-speech synthesis, emotional speech generation, speech expressiveness analysis, and speech emotion computing research tasks.
数据集概述:MOSS-Local Voice Acting v4 "emorant" (part 4/4)
名称: MOSS-Local Voice Acting v4 "emorant" — maximally emotional deliveries, 64 takes per group (part 4/4)
许可证: Apache-2.0
任务类别: 文本转语音 (text-to-speech)
语言: 英语
数据集用途与特点
该数据集是四部分语料库的第4部分,包含40种情绪 × 200组 × 64次录制(四部分合计目标为512,000个音频片段)。使用合并后的4.55B参数本地Transformer MOSS-TTS语音表演模型生成,原生48 kHz采样率,温度参数1.0,为原始模型输出,无参考音频(模型根据普通人说话者的人物描述自行发明声音)。
v4版本的目标是强度:从情绪原型提示研究中重新编写指令,推动输出远超中性朗读语音的表现力。
- 每种情绪映射到一种表演类别(如失控咆哮、痛苦、悲伤、喜悦、温柔、惊奇,以及醉酒/戏弄/狂喜的特殊类别),并附带该类别的特定语音机制语言(如撕裂、破裂、啜泣、呼喊、低语……)
- 情感名称采用发自内心的方式命名,指令构建一个升级弧线("每句话都更加恶毒")
- 脚本中嵌入括号标注的发声爆发提示
- 结尾的防护指令("……但每个词仍然清晰")保持可理解性,可通过
wer验证每次录制
脚本是全新的单一提示(从未在v2/v3中使用过),格式为小写人物描述 + 指令,均为英语。
本部分包含的情绪
- Pleasure_Ecstasy
- Pride
- Relief
- Sadness
- Sexual_Lust
- Shame
- Sourness
- Teasing
- Thankfulness_Gratitude
- Triumph
数据内容格式
每个 data/<Emotion>.tar 压缩包包含一种情绪的64次录制组(<Emotion>_<group>_v<take>.flac,48 kHz单声道FLAC格式),以及 scores.parquet 和 cells.jsonl(包含每组的确切文本、指令和生成元数据)。
scores.parquet 文件每行对应一次录制,包含以下列:
| 列名 | 描述 |
|---|---|
gid, emotion, ei_head, prompt_idx, seed |
组ID、目标情绪、用作目标的移情洞察头部、提示索引、录制索引 |
vocal_burst |
指令中包含的非语言爆发提示 |
wer, inv_wer, hyp |
与脚本相比的词错误率(Parakeet-TDT-0.6b-v3 ASR)、1/(1+WER)、ASR转录文本 |
blend, genu, prompt_sim |
VoiceCLAP-商业情绪融合分数、真实度分数、音频-文本与风格指令的余弦相似度 |
ei_* (42列) |
移情洞察-语音-Plus分数:40个情绪维度 + ei_Arousal、ei_Valence(以及 ei_Authenticity) |
rms_db, peak_db, dur, finished |
响度、峰值电平、时长(秒)、生成是否以EOS结束(而非达到token上限) |
target |
该组所提示情绪的EI-Plus分数 |
数据集未经过滤:所有录制均包含在内,用户可根据需要自行应用最佳N选择(例如每组使用 (norm(blend) + norm(genu)) * inv_wer)。
相关资源
- 生成模型:laion/moss-tts-local-transformer-4.55b-voice-acting(合并的4.55B参数本地Transformer MOSS-TTS,48 kHz)
- 推理代码与指南:LAION-AI/laion-moss-local-1.5-voice-acting-4.55b
- 姊妹发布:v2无参考语料库
laion/moss-local-voice-acting-64x100-part1..4,完整DramaBox重新诠释laion/moss-local-dramabox-full-reinterpretations-64




