davidguzmanr/open-bible-resources
收藏官方服务:
资源简介:
该数据集是一个多语言圣经音频-文本集合,涵盖多种语言配置(如Apali、Arabic Standard、Assamese等),每个配置包含音频特征和对应的文本转录,以及圣经相关的元数据(如testament、book、chapter、verse),还有持续时间(秒)和说话者ID。数据集分为训练和测试分割,用于语音识别、文本到语音或其他NLP任务。
This dataset is a multilingual Bible audio-text corpus covering multiple language configurations (e.g., Apali, Standard Arabic, Assamese, etc.). Each configuration includes audio features, corresponding text transcripts, Bible-related metadata such as testament, book, chapter, verse, as well as duration in seconds and speaker ID. The dataset is split into training and test splits, and is intended for speech recognition, text-to-speech (TTS) or other natural language processing (NLP) tasks.
提供机构:
davidguzmanr


