遇见数据集

Aynursusuz/tts-pretrain-refs-3k-mos

收藏
Hugging Face2026-04-27 更新2026-05-03 收录
官方服务:

资源简介:

tts-pretrain-refs-3k-mos数据集是一个用于文本转语音(TTS)任务的参考数据集,包含3000个训练样本。每个样本包含44.1 kHz单声道语音音频、对应的文本转录以及自动预测的声音质量分数(MOS)。MOS分数基于ITU-T MOS量表(1-5分),评分方法使用了Microsoft DNSMOS P.835 ONNX模型,并对原始模型输出进行了多项式后拟合。数据集的平均MOS分数为3.394。

The tts-pretrain-refs-3k-mos dataset is a reference dataset for text-to-speech (TTS) tasks, containing 3000 training samples. Each sample includes 44.1 kHz mono speech audio, corresponding text transcript, and an automatically predicted sound-quality score (MOS). The MOS score is based on the ITU-T MOS scale (1-5), and the scoring method uses the Microsoft DNSMOS P.835 ONNX model with a polynomial post-fit applied to the raw model outputs. The dataset has an average MOS score of 3.394.

提供机构:
Aynursusuz
二维码
社区交流群
二维码
科研交流群
商业服务