2M-BELEBELE
收藏资源简介:
2M-BELEBELE是一个高度多语言的语音和美式手语理解数据集,由Meta的FAIR团队扩展BELEBELE数据集创建。该数据集涵盖了74种口语和1种手语(美式手语),包含488个不同的段落、900个问题和每个问题的4个多选答案。数据集的创建过程包括人类录音和手语翻译,确保了数据的高质量和多样性。该数据集主要用于多模态理解任务,旨在解决低资源语言和手语的自然语言处理问题,提升语音和手语翻译的性能。
2M-BELEBELE is a highly multilingual speech and American Sign Language (ASL) understanding dataset, developed by extending the original BELEBELE dataset via Meta's FAIR team. This dataset encompasses 74 spoken languages and 1 sign language (American Sign Language), consisting of 488 distinct paragraphs, 900 questions, and 4 multiple-choice options for each question. The construction of the dataset involves human audio recording and sign language translation, which guarantees its high quality and diversity. Primarily designed for multimodal understanding tasks, this dataset aims to tackle natural language processing challenges faced by low-resource languages and sign languages, and enhance the performance of speech and sign language translation.




