humaneness-voice-dpo-rebalanced-six-task-v1
收藏资源简介:
该数据集名为“Humaneness Voice — Rebalanced Six-Task DPO Mix v1”,是一个用于语音合成偏好学习(DPO)的训练混合集。它包含总计447,179个训练偏好对和5,353个验证偏好对,以Zstandard压缩的Parquet分片文件存储。数据来源于六个任务家族:realspeech_stop(10万对)、voice_emotion(10万对)、voice_truncation(10万对)、cfg_p2_length(55,995对)、s5_pitch(66,805对)和quality_repair_sidon(24,379对)。每个样本包含chosen和rejected两种音频的MOSS Audio Tokenizer V2代码块(小端uint16数组,每80毫秒帧对应12个码本值),以及相关元数据(如家庭标签、对ID、源UID、语言、文本、提示等)。数据集主要用于训练文本转语音(TTS)系统中的偏好优化模型。注意:pitch任务存在严重的质量混淆,不应作为纯说话人身份数据集使用;验证集仅覆盖四个家族,缺少s5_pitch和quality_repair_sidon的验证对。所有音频由MOSS V2解码器生成,不包含原始波形。许可证为CC-BY-4.0。
The dataset is named "Humaneness Voice — Rebalanced Six-Task DPO Mix v1", a training mixture for preference learning (DPO) in speech synthesis. It contains a total of 447,179 training preference pairs and 5,353 validation preference pairs, stored as Zstandard-compressed Parquet shards. The data originates from six task families: realspeech_stop (100k pairs), voice_emotion (100k pairs), voice_truncation (100k pairs), cfg_p2_length (55,995 pairs), s5_pitch (66,805 pairs), and quality_repair_sidon (24,379 pairs). Each sample includes MOSS Audio Tokenizer V2 code blocks (little-endian uint16 arrays, 12 codebook values per 80ms frame) for both chosen and rejected audio, along with associated metadata (e.g., family tags, pair IDs, source UIDs, language, text, prompts). The dataset is primarily used for training preference optimization models in text-to-speech (TTS) systems. Note: The pitch task has severe quality confounding and should not be used as a pure speaker identity dataset; the validation set covers only four families, missing validation pairs for s5_pitch and quality_repair_sidon. All audio is generated by the MOSS V2 decoder, with no original waveforms included. License: CC-BY-4.0.




