QCRI/MMCQA-SemEval27
收藏资源简介:
MMCultureQA-SemEval27是一个用于SemEval 2027共享任务的多模态数据集,专注于基于文化的视觉问答,支持英语和阿拉伯语。系统接收一张图像和一个关于图像的问题(问题以音频或文本形式提供),并生成一个简短的开放式答案。问题常涉及食物、地点、习俗和物体,因此正确答案通常依赖于本地文化知识而非图像表面可见内容。数据集包含两个任务:Task 1(Spoken Visual QA)问题以口语音频片段形式给出;Task 2(Textual Visual QA)问题以文本形式给出。每个任务按语言变体分为不同轨道,目前包括英语和现代标准阿拉伯语,计划扩展至埃及阿拉伯语、黎凡特阿拉伯语等。数据文件包括图像、音频和JSONL文件,其中JSONL文件包含ID、图像路径、国家、类别、子类别、问题(或音频路径)和答案等字段。该版本为小样本发布,完整训练和开发集将后续发布。
MMCultureQA-SemEval27 is a multimodal dataset for the SemEval 2027 Shared Task, focusing on culture-based visual question answering. It supports both English and Arabic. The system takes an image and a question about the image (the question is provided in either audio or text format) and generates a short open-ended answer. Questions often relate to food, locations, customs, and objects, so correct answers typically rely on local cultural knowledge rather than visually apparent content in the image. The dataset includes two tasks: Task 1 (Spoken Visual QA), where questions are provided as spoken audio clips; Task 2 (Textual Visual QA), where questions are provided in text format. Each task is divided into distinct tracks based on language variants, currently covering English and Modern Standard Arabic, with plans to expand to Egyptian Arabic, Levantine Arabic, and other variants. The dataset files include images, audio files, and JSONL files. The JSONL files contain fields such as ID, image path, country, category, subcategory, question (or audio path), and answer. This release is a few-shot version, and the full training and development sets will be released subsequently.




