Saar-Voice
收藏资源简介:
Saar-Voice是由萨尔兰大学团队构建的德国萨尔布吕肯方言多说话人语音语料库,包含9名说话人录制的6小时方言语音数据。数据集通过数字化印刷书籍(66.6%)、本地社区文本(32.4%)及MASSIVE数据集本地化翻译(1%)三重来源构建,涵盖诗歌、散文、民间故事等文体,共8,772个句子75,280词。该语料库采用专业录音设备在隔音室采集,包含对齐的文本-音频表征,旨在解决德语方言文本转语音(TTS)任务中低资源方言数据缺失问题,为零样本和少样本模型适配提供研究基础。
Saar-Voice is a multi-speaker Saarbrücken German dialect speech corpus constructed by the team from Saarland University. It contains 6 hours of dialect speech data recorded by 9 speakers. The corpus is built from three sources: digitized printed books (66.6%), local community texts (32.4%), and localized translations from the MASSIVE dataset (1%). It covers various genres including poetry, prose, and folk tales, with a total of 8,772 sentences and 75,280 words. The corpus was collected in a soundproof room using professional recording equipment and includes aligned text-audio representations. It aims to address the shortage of low-resource dialect data for German dialect text-to-speech (TTS) tasks, providing a research foundation for zero-shot and few-shot model adaptation.

- 1Saar-Voice: A Multi-Speaker Saarbrücken Dialect Speech Corpus萨尔兰大学·语言科学与技术系 · 2026年



