遇见数据集

moss-local-dramabox-full-reinterpretations-64

收藏
魔搭社区2026-07-19 更新2026-07-19 收录
官方服务:

资源简介:

# MOSS-Local DramaBox Full-Performance Reinterpretations — best-of-64, dual-reward scored **2,953 full two-scene DramaBox voice-acting performances (EN/DE/ES/FR, all 9 DramaBox pathways), each reinterpreted 64 times** by [`laion/moss-tts-local-transformer-4.55b-voice-acting`](https://huggingface.co/laion/moss-tts-local-transformer-4.55b-voice-acting) (4.55B local transformer, native 48 kHz), **voice-cloned from the original performance**, with the complete scoring needed to select the best takes two different ways — voice identity or emotional performance. Source performances: [`TTS-AGI/dramabox-gemini-finetune`](https://huggingface.co/datasets/TTS-AGI/dramabox-gemini-finetune) (Gemini-prompted DramaBox two-scene performances, "CUT TO:" format). Methodology mirrors the [reproduce-and-improve study](https://projects.laion.ai/moss-8b-voice-acting/reproduce_study/index.html), scaled from best-of-8 to **best-of-64**. ## How each group was made - **instruction** = the sample's full Gemini DramaBox prompt (voice description + stage directions + quoted dialogue with vocal-burst notes, both scenes incl. "CUT TO:") - **text** = scene-1 + scene-2 expected texts - **reference** = the ORIGINAL full performance audio (both parts, in order) — encoded to 12-codebook MOSS v2 codes for voice cloning, and kept as the comparison target - generation: `audio_temperature 1.0` (with-reference setting), top-p 0.95, top-k 25, repetition 1.1, batched 64-at-once, token budget = max(words×6, ref_frames×1.2) at 12.5 Hz - audio: **raw native 48 kHz FLAC** — no post-processing ## Per-take scores (`scores.parquet`, also per-bucket inside each tar) | column | meaning | |---|---| | `wer`, `inv_wer` | word error rate vs the expected text ([`nvidia/parakeet-tdt-0.6b-v3`](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3), bf16 batched) | | `genu`, `blend` | genuineness & vocal-burst-blend heads on [`laion/voiceclap-commercial`](https://huggingface.co/laion/voiceclap-commercial) | | `prompt_sim` | VoiceCLAP cosine of direction-text vs audio | | `ei_*` (42 cols) | **Empathic-Insight-Voice-Plus** scores: 40 EmoNet emotions + Valence + Arousal ([`laion/Empathic-Insight-Voice-Plus`](https://huggingface.co/laion/Empathic-Insight-Voice-Plus) on [`laion/BUD-E-Whisper`](https://huggingface.co/laion/BUD-E-Whisper)) | | `emonet_42` | the same 42 values as one vector (list) | | `emotion_cos` | cosine(take's 42-vec, reference's 42-vec) — "same feeling?" | | `ecapa_192` | **ECAPA-TDNN speaker embedding** (192-d list, [`speechbrain/spkrec-ecapa-voxceleb`](https://huggingface.co/speechbrain/spkrec-ecapa-voxceleb)) | | `spk_sim` | cosine(take's ECAPA, reference's ECAPA) — "same voice?" | | `dur_match` | min(dur, ref_dur)/max(dur, ref_dur) | | `rms_db`, `peak_db`, `dur` | loudness & length | `references.parquet` holds each group's reference 42-dim EmoNet vector + ECAPA embedding. ## The two rewards (compute from the raw columns; min-max normalise globally) ``` reward_A (voice identity) = mean( n(spk_sim), n(inv_wer), n(genu), n(blend), n(dur_match) ) reward_B (emotion profile) = emotion_cos × inv_wer ``` Ranking A asks "same speaker, intelligible, at least as human as the original?"; Ranking B asks "does it feel like the same performance?". They disagree often — that is the point. ## Layout ``` data/bucket_XXX.tar ~100 groups each: <gid>/<gid>_vNN.flac (64 takes, 48 kHz) + bucket_XXX_scores.parquet scores.parquet all takes, all scores + embeddings references.parquet per-group reference vectors/embeddings ``` ## Generation stack Model [`laion/moss-tts-local-transformer-4.55b-voice-acting`](https://huggingface.co/laion/moss-tts-local-transformer-4.55b-voice-acting) · codec [`OpenMOSS-Team/MOSS-Audio-Tokenizer-v2`](https://huggingface.co/OpenMOSS-Team/MOSS-Audio-Tokenizer-v2) · recipe & code: [github.com/LAION-AI/laion-moss-local-1.5-voice-acting-4.55b](https://github.com/LAION-AI/laion-moss-local-1.5-voice-acting-4.55b) · 8×A100, fused generate→score workers (~1.3–1.6 s/clip/GPU incl. all scoring). ## License **Apache-2.0.** Synthetic speech generated by the model named above; source prompts/performances from `TTS-AGI/dramabox-gemini-finetune` (CC-BY-4.0).

提供机构:
maas
创建时间:
2026-07-15
二维码
社区交流群
二维码
科研交流群
商业服务