moss-local-voice-acting-v4-emorant-64x200-part1
收藏资源简介:
MOSS-Local Voice Acting v4 "emorant"(第1部分/共4部分)是一个专注于最大化情感表达的语音合成数据集。该数据集使用合并后的4.55B本地变换器MOSS-TTS语音表演模型生成,采样率为48kHz,温度参数为1.0,采用原始模型输出,且无参考音频(模型根据普通人类说话者的人物描述创造声音)。数据集通过改写提示指令,基于情感原型研究,将语音表达推向远超中性朗读的水平。每个情感被映射到特定的表达类别(如无约束的咆哮、痛苦、悲伤、喜悦、温柔、惊奇等),并包含类别特定的声音机制描述(如撕裂声、破音、啜泣、欢呼、低语等)。提示采用感性描述语言,构建情感升级弧线,并在脚本中嵌入括号内的声音爆发提示,同时通过结尾保护机制保持可理解性。本部分数据集包含10种情感类型:喜爱、娱乐、愤怒、惊讶/惊奇、敬畏、苦涩、专注、困惑、沉思、轻蔑。每个情感对应一个.tar压缩文件,包含64个版本的音频组(FLAC格式)以及scores.parquet和cells.jsonl元数据文件。scores.parquet为每个音频版本提供详细评分数据,包括词错误率、情感混合分数、真实性分数、42维共情洞察语音评分以及音频特征(响度、时长、生成完成状态)等。数据集包含所有生成版本,未经过滤,用户可应用自定义的最佳选择策略。该数据集适用于文本到语音合成、情感语音生成、语音情感分析等任务,特别适合需要高情感表现力的语音合成研究。
MOSS-Local Voice Acting v4 "emorant" (Part 1 of 4) is a speech synthesis dataset focused on maximizing emotional expression. It is generated using a merged 4.55B local transformer MOSS-TTS voice acting model, with a sampling rate of 48kHz, a temperature parameter of 1.0, using raw model output, and no reference audio (the model creates voices based on character descriptions of ordinary human speakers). The dataset aims to achieve maximum intensity in emotional expression: by rewriting prompt instructions, based on emotional prototype research, it pushes speech expression far beyond neutral reading levels. Each emotion is mapped to specific expression categories (such as unrestrained roar, pain, sadness, joy, tenderness, surprise, etc.) and includes category-specific vocal mechanism descriptions (e.g., tearing sounds, vocal breaks, sobbing, cheering, whispering, etc.). The prompts use emotional descriptive language, construct emotional escalation arcs, and embed voice burst cues in parentheses within the script, while maintaining intelligibility through end protection mechanisms. This part of the dataset includes 10 emotion types: affection, amusement, anger, surprise/astonishment, awe, bitterness, concentration, confusion, contemplation, contempt. Each emotion corresponds to a .tar compressed file containing 64 versions of audio groups (in FLAC format) and metadata files scores.parquet and cells.jsonl. scores.parquet provides detailed scoring data for each audio version, including word error rate, emotional blend score, authenticity score, 42-dimensional empathetic voice insights score, and audio features (loudness, duration, generation completion status), etc. The dataset includes all generated versions, unfiltered, allowing users to apply custom best-selection strategies. It is suitable for text-to-speech synthesis, emotional speech generation, speech emotion analysis, and other tasks, particularly for speech synthesis research requiring high emotional expressiveness.
数据集概述:MOSS-Local Voice Acting v4 "emorant" (part 1/4)
- 数据集名称: MOSS-Local Voice Acting v4 "emorant" (part 1/4)
- 许可证: Apache-2.0
- 任务类别: 文本转语音(Text-to-Speech)
- 语言: 英语
- 数据集规模: 这是4部分语料库的第1部分,包含 40种情绪 × 200个组 × 64个录音,目标是在4个部分中生成512,000个音频片段。
- 生成模型: 使用合并的4.55B参数本地Transformer MOSS-TTS语音表演模型,以原生48 kHz采样率、温度1.0生成,无参考音频(模型根据普通人说话者的人物描述创造声音)。
- 核心特点: 强调情感强度。指令基于情感原型提示研究重新编写,推动表达远超中性朗读语音:
- 每种情绪映射到一个表达类别(如失控咆哮、痛苦、悲伤、喜悦、温柔、惊奇,以及醉酒/挑逗/狂喜等特殊类别),并附带类别特定的语音机制语言(如撕裂、破音、啜泣、欢呼、低语...)。
- 情感被具象化命名,指令构建升级弧线(例如"每句台词更加恶毒")。
- 脚本中嵌入了括号内的声音爆发提示。
- 结尾有保护性指令("...但每个词仍然清晰可辨"),可通过
wer(词错误率)验证可懂度。
- 脚本: 均为全新的单一提示(从未在v2/v3中使用过),格式为小写人物描述 + 指令,使用英语。
本部分包含的情绪
- Affection(喜爱)
- Amusement(娱乐)
- Anger(愤怒)
- Astonishment_Surprise(惊讶)
- Awe(敬畏)
- Bitterness(苦涩)
- Concentration(专注)
- Confusion(困惑)
- Contemplation(沉思)
- Contempt(轻蔑)
数据集内容
每个 data/<Emotion>.tar 文件包含一个情绪的64个录音组(文件格式为 <Emotion>_<group>_v<take>.flac,48 kHz单声道FLAC),以及 scores.parquet 和 cells.jsonl(包含每组的精确文本、指令和生成元数据)。
scores.parquet 包含每一项录音的详细指标,主要列如下:
| 列名 | 描述 |
|---|---|
gid, emotion, ei_head, prompt_idx, seed |
组ID、目标情绪、用作目标的Empathic-Insight头部、提示索引、录音索引 |
vocal_burst |
指令中包含的非语言爆发提示 |
wer, inv_wer, hyp |
与脚本的词错误率(Parakeet-TDT-0.6b-v3 ASR识别)、1/(1+WER) 和ASR转录文本 |
blend, genu, prompt_sim |
VoiceCLAP-commercial情绪混合分数、真实度分数、音频-文本与风格指令的余弦相似度 |
ei_*(42列) |
Empathic-Insight-Voice-Plus分数:40个情绪维度 + ei_Arousal(唤醒度)、ei_Valence(效价)和 ei_Authenticity(真实度) |
rms_db, peak_db, dur, finished |
响度(dB)、峰值(dB)、持续时间(秒)、以及生成是否以EOS(而非token上限)结束 |
target |
为该组提示的情绪的EI-Plus分数 |
所有录音都包含在内,未经过滤,用户可以根据自己的最佳N选策略(例如,按组计算 (norm(blend) + norm(genu)) * inv_wer)进行筛选。
相关资源
- 生成模型: laion/moss-tts-local-transformer-4.55b-voice-acting(合并的4.55B参数本地Transformer MOSS-TTS,48 kHz)
- 推理代码和指南: LAION-AI/laion-moss-local-1.5-voice-acting-4.55b
- 系列其他数据集: v2无参考语料库
laion/moss-local-voice-acting-64x100-part1..4,以及完整DramaBox重解读数据集 laion/moss-local-dramabox-full-reinterpretations-64
许可证和来源说明
- 许可证: Apache-2.0
- 音频均为全合成生成(无参考集不包含真实说话人的声音;v3的参考片段来自公开的
TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave数据集)




