disco-eth/EuroSpeech-24kHz
收藏资源简介:
EuroSpeech是一个大规模多语言语音语料库,包含22种欧洲语言的高质量对齐议会语音数据。该数据集通过处理议会会议记录构建,使用了鲁棒的对齐流程来处理不同的音频格式和非逐字转录。数据集包含约78,100小时的对齐语音文本数据,并提供了不同质量等级的子集(CER < 30%、< 20%、< 10%)。主要子集(CER < 20%)通过Hugging Face Datasets接口直接提供,适用于自动语音识别(ASR)、文本到语音(TTS)系统、多语言语音研究、低资源语言语音技术开发和跨语言迁移学习等用途。
EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. It includes approximately 78,100 hours of initially aligned speech-text data, with quality-filtered subsets (CER < 30%, < 20%, < 10%). The primary subset (CER < 20%) is provided directly through the Hugging Face Datasets interface for all languages, intended for automatic speech recognition (ASR), text-to-speech (TTS) systems, multilingual speech research, low-resource language speech technology development, and cross-lingual transfer learning in speech models.




