LoReSpeech
收藏资源简介:
LoReSpeech是一个针对低资源语言的语音对语音平行语料库,由奥尔良大学和Yanantic AI合作开发。该数据集旨在为低资源语言提供高质量的语音对齐资源,包括语音识别子集LoReASR和长篇语音记录的对齐。它通过结合精确的本地合作,生产高质量的数据,并直接涉及相关社区,以促进多语种语音识别系统、直接语音翻译模型、跨语言语言分析和濒危语言保护等领域的发展。
LoReSpeech is a speech-to-speech parallel corpus for low-resource languages, co-developed by the University of Orléans and Yanantic AI. This dataset aims to provide high-quality speech-aligned resources for low-resource languages, including the LoReASR subset for speech recognition and aligned long-form speech recordings. It is built through precise local collaboration and direct engagement with relevant communities to advance developments in fields including multilingual speech recognition systems, direct speech translation models, cross-linguistic language analysis, and endangered language conservation.
- 1Connecting Voices: LoReSpeech as a Low-Resource Speech Parallel Corpus法国奥尔良大学, Yanantic AI · 2025年



