OpenBibleTTS
收藏资源简介:
OpenBibleTTS是一个面向低资源语言的语音合成大规模基准数据集,由多个研究机构联合创建,旨在解决全球语言资源分布不均的问题。该数据集涵盖37种代表性不足的语言,包含总计约112万条语音-文本对齐数据,总时长超过3,469小时,数据来源于开放授权的圣经语音平台。其构建过程通过系统化的文本解析、音频分割和经文级对齐技术,将原始资源转化为可直接用于训练和评估的标准化语料。该数据集主要应用于低资源语音合成研究领域,为评估不同TTS架构和大规模语音生成模型在跨语言场景下的性能提供了重要基准,推动更具包容性的全球语音技术发展。
OpenBibleTTS is a large-scale benchmark dataset for speech synthesis in low-resource languages, jointly developed by multiple research institutions to address the global inequity in language resource distribution. This dataset encompasses 37 underrepresented languages, totaling approximately 1.12 million speech-text aligned pairs with a combined duration of over 3,469 hours. The data is sourced from open-licensed Bible speech platforms. During its construction, original resources are converted into standardized corpora directly applicable for training and evaluation via systematic text parsing, audio segmentation, and verse-level alignment techniques. Primarily utilized in low-resource speech synthesis research, this dataset provides a critical benchmark for evaluating the performance of diverse TTS architectures and large-scale speech generation models in cross-lingual scenarios, thereby advancing the development of more inclusive global speech technologies.

- 1OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages麦吉尔大学; 米拉魁北克人工智能研究所; AIMS研究与创新中心; NM-AIST; 萨尔大学; 加拿大CIFAR人工智能主席 · 2026年



