IndicVoices-R
收藏资源简介:
IndicVoices-R是由印度理工学院马德拉斯分校计算机科学与工程系创建的,是目前最大的多语言多说话者印度TTS数据集,涵盖了22种印度语言,包含1,704小时的高质量语音数据,来自10,496名说话者。数据集主要由即兴录音组成,确保了语音的自然性。创建过程中,通过去噪和语音增强模型对原始ASR数据进行处理,以提高语音质量。该数据集旨在解决印度语言TTS系统中数据稀缺和高多样性需求的问题,支持零样本、少样本和多样本说话者泛化能力的评估。
IndicVoices-R was developed by the Department of Computer Science and Engineering, Indian Institute of Technology Madras. It is currently the largest multilingual, multi-speaker Indian TTS dataset, covering 22 Indian languages and containing 1,704 hours of high-quality speech data from 10,496 speakers. The dataset primarily comprises spontaneous recordings to ensure speech naturalness. During its development, raw ASR data was processed via denoising and speech enhancement models to enhance audio quality. This dataset is designed to address the issues of data scarcity and high diversity requirements in Indian language TTS systems, and supports evaluation of zero-shot, few-shot, and multi-sample speaker generalization capabilities.
IndicVoices-R: 大规模多语言多说话人语音数据集,用于扩展印度TTS
摘要
IndicVoices-R (IV-R) 是基于ASR数据集生成的最大规模多语言印度TTS数据集,包含1,704小时的高质量语音数据,涵盖22种印度语言,来自10,496名说话人。IV-R数据集的质量与LJSpeech、LibriTTS和IndicTTS等黄金标准TTS数据集相当。此外,IV-R引入了IV-R基准,用于评估TTS模型在印度语音上的零样本、少样本和多样本说话人泛化能力,确保年龄、性别和风格的多样性。
资源
数据集下载地址:https://ai4bharat.iitm.ac.in/indicvoices_r/
清单格式
filename: 指向wav文件的路径text: 音频的转录文本,使用标准化版本duration: 音频时长(秒)lang: 语言的ISO代码samples: 样本数量verbatim: 转录文本的逐字版本normalized: 转录文本的标准化版本speaker_id: 唯一的说话人IDscenario: 数据类型task_name: 任务名称gender: 说话人性别age_group: 说话人年龄组job_type: 说话人职业类型qualification: 说话人学历area: 说话人所属地区district: 说话人所属区state: 说话人所属州occupation: 说话人职业verification_report: 由QA团队提供的验证标记chunk_name: 音频块名称snr: 信噪比c50: C50值utterance_pitch_mean: 语音音调均值utterance_pitch_std: 语音音调标准差cer: 字符错误率
许可证




