FMSD-TTS
收藏资源简介:
FMSD-TTS数据集是由电子科技大学信息与软件工程学院、西藏大学信息科学技术学院和德克萨斯大学西南医学中心眼科学系合作生成的,旨在解决藏语资源匮乏的问题。该数据集包含超过210小时的录音,涵盖了藏语三大主要方言——卫藏、安多和康巴,共计1,500多位母语者的音频样本,数据集大小为120,000条。数据集的生成过程中采用了FMSD-TTS模型,该模型能够从有限的参考音频和显式方言标签中合成平行方言语音。数据集的创建过程采用了先进的技术手段,包括讲者-方言融合模块和方言专用动态路由网络(DSDR-Net),能够捕捉不同方言之间的细微声学和语言变化,同时保持讲者身份。FMSD-TTS数据集的发布为藏语语音处理领域提供了宝贵的新资源,有助于推动自动语音识别(ASR)、语音翻译(ST)和语音-语音方言转换(S2SDC)等领域的研究。
The FMSD-TTS dataset was collaboratively developed by the School of Information and Software Engineering of the University of Electronic Science and Technology of China, the School of Information Science and Technology of Tibet University, and the Department of Ophthalmology of the University of Texas Southwestern Medical Center, aiming to address the shortage of Tibetan language resources. This dataset contains over 210 hours of recordings covering the three major Tibetan dialects: Ü-Tsang, Amdo, and Khams, with audio samples from more than 1,500 native speakers, totaling 120,000 entries. The dataset was generated using the FMSD-TTS model, which can synthesize parallel dialectal speech from limited reference audio and explicit dialect labels. The dataset construction adopted advanced technical approaches, including a speaker-dialect fusion module and a dialect-specific dynamic routing network (DSDR-Net), which can capture subtle acoustic and linguistic variations across different dialects while preserving speaker identity. The release of the FMSD-TTS dataset provides a valuable new resource for the field of Tibetan speech processing, and helps advance research in areas including automatic speech recognition (ASR), speech translation (ST), and speech-to-speech dialect conversion (S2SDC).
数据集概述
基本信息
- 标题: FMSD-TTS: Few-shot Multi-Speaker Multi-Dialect Text-to-Speech Synthesis for Ü-Tsang, Amdo and Kham Speech Dataset Generation
- arXiv标识符: arXiv:2505.14351v1
- 提交日期: 2025年5月20日
- 领域: 计算机科学 > 语音 (cs.SD)
- 作者: Yutong Liu, Ziyue Zhang, Ban Ma-bao, Yuqing Cai, Yongbin Yu, Renzeng Duojie, Xiangxiang Wang, Fan Gao, Cheng Huang, Nyima Tashi
摘要
- 研究背景: 藏语是一种低资源语言,其三大主要方言(Ü-Tsang、Amdo和Kham)的平行语音语料库稀缺,限制了语音建模的进展。
- 解决方案: 提出FMSD-TTS,一种少样本、多说话人、多方言的文本到语音合成框架,能够从有限的参考音频和明确的方言标签中合成平行方言语音。
- 创新点:
- 新颖的说话人-方言融合模块。
- 方言专用动态路由网络(DSDR-Net),用于捕捉跨方言的细粒度声学和语言变化,同时保留说话人身份。
- 评估: 通过客观和主观评估,FMSD-TTS在方言表达和说话人相似性方面显著优于基线。
- 贡献:
- 专为藏语多方言语音合成设计的少样本TTS系统。
- 公开发布由FMSD-TTS生成的大规模合成藏语语音语料库。
- 开源评估工具包,用于标准化评估说话人相似性、方言一致性和音频质量。
技术细节
- 评论: 13页
- 主题分类:
- 语音 (cs.SD)
- 人工智能 (cs.AI)
- 计算与语言 (cs.CL)
- 音频与语音处理 (eess.AS)
- DOI: 10.48550/arXiv.2505.14351
相关资源
- 全文链接:




