QCRI/OASIS
收藏资源简介:
OASIS是一个大规模的多语言和多模态数据集,专注于文化和视觉问答。它旨在评估多模态模型在对象识别之外的能力,特别是在现实场景中的实用、常识和文化基础推理。数据集包含约92万张真实图像、1480万个问答对、370万个语音问题、383小时的人类录音和2万小时的语音克隆数据,覆盖英语和阿拉伯语的18个国家的多种变体。OASIS支持文本、语音、图像及其组合的多种输入设置,适用于多语言和多模态问答、文化基础推理等研究。
OASIS is a large-scale culturally grounded multimodal question answering dataset covering images, text, and speech. It is designed to evaluate multimodal models beyond object recognition, with emphasis on pragmatic, commonsense, and culturally grounded reasoning in real-world scenarios. The dataset contains approximately 0.92M real images, 14.8M QA pairs, 3.7M spoken questions, 383 hours of human-recorded speech, and 20K hours of voice-cloned speech, covering English and Arabic varieties across 18 countries. OASIS supports four input settings: text-only, speech-only, text + image, and speech + image, and is intended for research on multimodal and multilingual question answering, culturally grounded reasoning, and more.




