laion-voice-profiles-dpo-cfg
收藏资源简介:
该数据集是LAION Voice Profiles项目的一部分,名为“LAION Voice Profiles — contrastive DPO pairs (CFG + phase 2)”,由Christoph Schuhmann和LAION团队创建。它包含842,935个偏好对(preference pairs),分为四个家族:cfg_high(237,209对)、cfg_low(237,209对)、p2_emox(309,128对)和p2_len(59,389对)。这些数据基于与`laion/laion-voice-profiles-sft`和`laion/laion-voice-profiles-dpo`相同的500个合成语音profile,但专注于强度对比和选择性,以解决之前模型中的缺陷。数据集以两个配置(config)形式提供:`cfg`(包含cfg_high和cfg_low)和`p2`(包含p2_emox和p2_len)。每个配置包含128个parquet分片(shard),采用`zstd`压缩,共计31列。所有样本均为机器生成,语言为英语和德语。数据集的列包括:`uid`(唯一对ID)、`family`(家族名称)、`src`(源组,均为0,表示合成语音profile)、`voice_key`(语音profile ID,500个之一)、`lang`(所选剪辑的语言)、`text`(所选剪辑的生成文本)、`caption_general`(强度陈述)、`frames`(MOSS帧数)、`ref_codes`(参考剪辑的MOSS代码)、`codes`(所选剪辑代码)、`rej_codes`(被拒绝剪辑代码)等。该数据集适用于文本转语音(TTS)任务中的偏好优化、DPO训练、语音克隆、语音情感控制等场景。
This dataset is part of the LAION Voice Profiles project, named LAION Voice Profiles — contrastive DPO pairs (CFG + phase 2), created by Christoph Schuhmann and the LAION team. It contains 842,935 preference pairs, divided into four families: cfg_high (237,209 pairs), cfg_low (237,209 pairs), p2_emox (309,128 pairs), and p2_len (59,389 pairs). The data is based on the same 500 synthetic voice profiles as laion/laion-voice-profiles-sft and laion/laion-voice-profiles-dpo, but focuses on intensity contrast and selectivity to address previous model deficiencies. The dataset is provided in two configurations: cfg (including cfg_high and cfg_low) and p2 (including p2_emox and p2_len). Each configuration contains 128 parquet shards compressed with zstd, totaling 31 columns. All samples are machine-generated, and the languages are English and German. Columns include: uid (unique pair ID), family (family name), src (source group, always 0 for synthetic voice profiles), voice_key (voice profile ID, one of 500), lang (language of the selected clip), text (generated text of the selected clip), caption_general (intensity statement), frames (MOSS frame count), ref_codes (MOSS codes of reference clip), codes (selected clip codes), rej_codes (rejected clip codes), etc. The dataset is suitable for preference optimization, DPO training, voice cloning, voice emotion control, and other tasks in text-to-speech (TTS) scenarios.




