西班牙语自动评估TTS质量的数据库
收藏资源简介:
本研究开发了一个用于自动评估西班牙语文本语音转换系统(TTS)质量的数据库,旨在提高自然度预测模型的准确性。该数据集包含来自52个不同的TTS系统和人声的4,326个音频样本,据我们所知,这是西班牙语中第一个此类数据集。为了对音频进行标注,设计了一个基于ITU-T Rec. P.807标准的客观测试,并由92名参与者完成。此外,通过训练自动自然度预测系统验证了收集到的数据集的实用性。我们探索了两种方法:在为英语训练的现有模型上进行微调,以及在冻结的自监督语音模型之上训练小型下游网络。我们的模型在五分制的MOS尺度上实现了0.8的平均绝对误差。进一步的分析表明了开发的数据库的质量和多样性,以及其在西班牙语TTS研究中的潜在价值。
This study developed a database for automatically evaluating the quality of Spanish text-to-speech (TTS) systems, aiming to improve the accuracy of naturalness prediction models. The dataset contains 4,326 audio samples from 52 distinct TTS systems and human voices, and, to the best of our knowledge, this is the first such dataset in Spanish. To annotate the audio, an objective test based on the ITU-T Rec. P.807 standard was designed and completed by 92 participants. Additionally, the utility of the collected dataset was verified by training automatic naturalness prediction systems. We explored two approaches: fine-tuning an existing model trained for English, and training a small downstream network on top of a frozen self-supervised speech model. Our model achieved a Mean Absolute Error (MAE) of 0.8 on the 5-point Mean Opinion Score (MOS) scale. Further analysis demonstrates the quality and diversity of the developed database, as well as its potential value in Spanish TTS research.
数据集概述
基本信息
- 数据集名称:es-TTS-subjective-naturalness
- 研究论文:INTERSPEECH 2025论文《A Dataset for Automatic Assessment of TTS Quality in Spanish》
- 数据集地址:https://huggingface.co/datasets/asosawelford/es-TTS-subjective-naturalness
数据集内容
- 用途:用于西班牙语TTS(文本到语音)质量自动评估
- 评估维度:主观自然度(subjective naturalness)
相关资源
- 源代码及附加结果:包含在TTS_dataset_analysis仓库中




