humaneness-voice-small-dpo-acting-eval
收藏资源简介:
Humaneness Voice Small DPO 匹配表演评估音频数据集是一个用于文本到语音(TTS)模型评估的音频集合。该数据集提供了 S3 基础模型与其四个 S3 DPO LoRA 适配器之间的可听五路比较。每个适配器均在固定的 Humaneness Voice Small acting benchmark 上进行了评估,该基准包含 50 个场景、11 种提示条件,涵盖英语和德语,使用相同的脚本、提示、参考分配和生成种子。每个模型有 1,872 个匹配片段,四个 LoRA 适配器贡献了 7,488 个可直接访问的 128 kbps 单声道 MP3 文件。数据集的比较键为 (challenge_id, surface, mode, seed),MP3 文件路径结构清晰。此外,数据集还包含 scores 目录下的 16 个 Parquet 文件和 summary.json,提供 ASR/WER、真实性、情感向量等评估指标,以及 proofs 目录中的不可变记录。该数据集适用于 TTS 模型的性能比较、DPO 微调效果分析以及自然度评估,但需注意预测指标存在局限性,应结合人工试听进行判断。
Humaneness Voice Small DPO matched performance evaluation audio dataset is an audio collection for evaluating text-to-speech (TTS) models. It provides an audible five-way comparison between the S3 base model and its four S3 DPO LoRA adapters. Each adapter is evaluated on the fixed Humaneness Voice Small acting benchmark, which includes 50 scenes, 11 prompt conditions, covering English and German, using the same scripts, prompts, reference assignments, and generation seeds. Each model has 1,872 matched segments, and the four LoRA adapters contribute 7,488 directly accessible 128 kbps mono MP3 files. The comparison key of the dataset is (challenge_id, surface, mode, seed), and the MP3 file path structure is clear. Additionally, the dataset includes 16 Parquet files and summary.json in the scores directory, providing evaluation metrics such as ASR/WER, naturalness, emotion vectors, as well as immutable records in the proofs directory. This dataset is suitable for TTS model performance comparison, DPO fine-tuning effect analysis, and naturalness evaluation, but it should be noted that the prediction metrics have limitations and should be combined with human listening for judgment.
Humaneness Voice Small DPO: matched acting evaluation audio
数据集概述
- 这是一个匹配的、可聆听的五方对比数据集。
- 包含原始的 S3 基础模型,以及四个 S3 DPO LoRA 适配器(原始/重新平衡的任务混合 × rank 64/128)。
- 每个 LoRA 都在与 S3 相同的、冻结的 Humaneness Voice Small acting benchmark 上进行了评测。
- 评测基准包含 50 个场景、11 种提示条件、英语和德语,并使用相同的脚本、提示、参考分配和生成种子。
- 每个模型有 1,872 条匹配的录制样本。
- 四个 LoRA 在此数据集中贡献了 7,488 个可直接访问的 128-kbps 单声道 MP3 文件;S3 的 MP3 仍保留在现有的基准仓库中,并通过并排聆听 Space 进行链接。
对比键与文件路径
- 对比键为
(challenge_id, surface, mode, seed)。 - 适配器的 MP3 路径为
mp3/{model}/{challenge_id}__{surface}__{mode}__seed{seed}.mp3。 - 例如:
mp3/original_rank064/HVSAC-001__C01__instruction__seed17.mp3。 - 每个 MP3 都可以通过 Hub 的
resolve/main/...URL 配合浏览器范围请求进行播放。
数据文件结构
scores/{model}/rankNNN.parquet:共 16 个文件,每个文件保留一条录制样本对应的一行数据,包含完整的 Parakeet ASR/WER、genuineness、vocal-burst blending、Empathic Insight Voice Plus 40-emotion vector、VoiceNet vector、clip-duration diagnostics、prompt 和 provenance。proofs/:包含不可变的按 rank 的 PASS 记录和文件哈希。summary.json:报告匹配行统计信息和缺失情况。
如何解读结果
- Space 仅展示公共键匹配的对比,而非不相关的随机样本。
- WER 越低越好。
- genuineness 和 blend predictor 值越高通常表示模型预测的自然度/混合度更强,但预测器并不完美,不能替代聆听。
- 目标情感激活是该场景所请求情感头的平均输出;它是一种预测器激活值,而非校准的情感准确率百分比。
- Vocal-burst F1 仅在请求 bursts 的场景上进行汇总。
- Clip 级时长误差仅作为诊断用途。
- 强制句子对齐尚未完成,因此本次发布不声称已验证逐句时序合规性。
- 即使生成和首轮评分成功,源 PASS 文件也可能带有
publication_ready: false,以表示缺失对齐后处理步骤。 - 在选择适配器之前,请自行比较提示和音频。
训练与许可信息
- 四个适配器检查点均源自同一 S3 基础模型。
- 原始 DPO 配对混合包含 715,402 对;重新平衡的六任务混合包含 447,179 对。
- 重新平衡的运行没有独立的配对级音高或质量留出集,因此配对准确率无法证明这两项新增内容能改善音频。
- 完整的训练注意事项和可复现代码参见模型卡。
- 底层的 acting-challenge prompts 来自
laion/acting-challenge-dataset,采用 Apache 2.0 许可。 - 生成的音频及本衍生评测发布采用 CC BY 4.0 许可,与 S3 基础模型一致。
- 请勿将此基准用作说话人同意或身份保真度的测试。
元数据
- 许可证:cc-by-4.0
- 语言:英语(en)、德语(de)
- 任务类别:text-to-speech
- 标签:audio、evaluation、humaneness-voice-small、dpo




