MaSS
收藏资源简介:
MaSS数据集是由法国格勒诺布尔-阿尔卑斯大学的研究团队创建的一个大规模、干净的多语言句子对齐口语语料库。该数据集基于圣经文本,涵盖了8种语言(巴斯克语、英语、芬兰语、法语、匈牙利语、罗马尼亚语、俄语和西班牙语),共包含8160个平行口语语句,适用于多种语言学研究,如语音对齐和翻译。数据集的创建过程涉及自动化的语音到文本和语音到语音的对齐技术,并通过人工评估确保了数据质量。MaSS数据集的应用领域广泛,包括自动语音识别、语音到语音翻译和语音检索等,旨在解决多语言环境下的语音处理问题。
The MaSS dataset is a large-scale, clean multilingual sentence-aligned spoken corpus created by a research team from Université Grenoble Alpes, France. Based on biblical texts, it covers 8 languages: Basque, English, Finnish, French, Hungarian, Romanian, Russian and Spanish, and contains a total of 8,160 parallel spoken utterances. It is suitable for various linguistic research topics such as speech alignment and translation. The dataset construction process involves automated speech-to-text and speech-to-speech alignment techniques, and its data quality is ensured through manual evaluation. The MaSS dataset has a wide range of application areas, including automatic speech recognition, speech-to-speech translation, speech retrieval and others, aiming to address speech processing issues in multilingual environments.



