meetween/mumospee_v1_fix
收藏资源简介:
Mumospee是一个全面的多语言语音元数据语料库,支持Meetween项目实现跨虚拟环境的包容性、语言无关协作。该数据集提供了来自公开可用数据集的语音音频元数据和下载URL的精选集合,优化了高性能计算集群的处理。它包括140,084小时的语音元数据,涵盖53,983,241个样本,覆盖25种欧盟语言及其他语言,平均每个样本时长9.34秒,平均转录长度16.5个单词。数据集结构包含音频路径、URL、类型、时长、语言、转录文本、标签(如CoVoST、GigaSpeech、PeopleSpeech、Librispeech、LibriTTS、Emilia、MOSEL)、分割(训练、测试、验证)和许可证字段。数据集旨在支持SpeechLLM和其他大型语言模型,以促进语言无关的虚拟会议应用,数据来源于CoVoST、GigaSpeech、PeopleSpeech、LibriSpeech、LibriTTS、Emilia、MOSEL等公开数据集。
Mumospee is a comprehensive multilingual speech-metadata corpus that supports the Meetween projects mission of enabling inclusive, language-neutral collaboration across virtual environments. The dataset provides metadata and download URLs for a curated collection of speech audio sourced from publicly available datasets, optimized for processing on high-performance computing clusters. It features 140,084 hours of speech metadata across 53,983,241 samples, covering 25 EU languages plus additional languages, with an average duration of 9.34 seconds per sample and an average transcript length of 16.5 words. The dataset structure includes fields such as path, URL, type, duration, language, transcript, tag (e.g., CoVoST, GigaSpeech, PeopleSpeech, Librispeech, LibriTTS, Emilia, MOSEL), split (train, test, validation), and license. It is designed to enable SpeechLLM and other large language models to support language-neutral virtual meeting applications, with data sourced from publicly available datasets like CoVoST, GigaSpeech, PeopleSpeech, LibriSpeech, LibriTTS, Emilia, and MOSEL.




