Fleurs-SLU
收藏资源简介:
Fleurs-SLU是一个大规模多语言口语理解基准,由维尔茨堡大学、剑桥大学和Mila研究所共同创建。该数据集包含102种语言的主题语音分类任务和92种语言的听力理解多选问答任务,数据来源于Fleurs、Flores、SIB-200和Belebele等数据集。数据集创建过程中,首先从Fleurs中过滤掉静音和噪声实例,然后将其与Flores、SIB-200和Belebele进行对齐和合并。Fleurs-SLU旨在解决低资源语言的语音识别和理解问题,特别是在缺乏正式书写系统的语言中,提升多语言语音技术的鲁棒性和包容性。
Fleurs-SLU is a large-scale multilingual spoken language understanding benchmark jointly developed by the University of Würzburg, the University of Cambridge, and Mila. This dataset includes topic speech classification tasks across 102 languages and multiple-choice listening comprehension question answering tasks for 92 languages, with data sourced from existing datasets such as Fleurs, Flores, SIB-200, and Belebele. During the dataset construction process, silent and noisy instances were first filtered out from Fleurs, followed by alignment and merging with Flores, SIB-200, and Belebele. Fleurs-SLU aims to address speech recognition and understanding challenges for low-resource languages, particularly those without a formal writing system, to improve the robustness and inclusivity of multilingual speech technologies.




