QCRI/AynVQA-ArabicNLP26
收藏资源简介:
Ayn-VQA-ArabicNLP26 是一个基于阿拉伯文化的多模态评估数据集,属于ArabicNLP 2026会议中的ImageEval 2026共享任务的一部分。该数据集旨在测试模型是否能够理解文化特定的图像,通过阿拉伯语口语问题以及区分真实描述与幻觉描述。数据集包含两个任务:口语视觉问答(给定图像和口语问题及选项音频,选择正确选项)和幻觉检测(给定图像和三个陈述,判断每个陈述是否为真实或幻觉,其中仅有一个陈述是真实的)。每个任务提供英语和现代标准阿拉伯语(MSA)两个语言版本,且这两个版本在图像和答案上并行,问题互为翻译。数据集覆盖18个阿拉伯国家,包括阿尔及利亚、巴林、埃及等。音频部分在训练、开发和开发测试集中使用合成语音生成,而最终测试集将使用人工录制音频。数据集文件包括图像、音频和JSONL文件,其中JSONL文件包含任务相关的字段,如ID、图像路径、音频路径和标签。数据划分包括训练集(3000项,带标签)、开发集(500项,带标签)、开发测试集(500项,无标签)和测试集(1000项,无标签)。评估指标针对每个任务单独设计,口语视觉问答使用准确率作为排名指标,幻觉检测使用对比不稳定性(CI)作为排名指标。数据集还提供了基线模型和示例笔记本,许可证为CC BY-NC-SA 4.0。
Ayn-VQA-ArabicNLP26 is a culturally grounded Arabic multimodal evaluation dataset, part of the ImageEval 2026 Shared Task at ArabicNLP 2026. It tests whether a model can read a culturally specific image, both from a spoken Arabic question and by telling grounded descriptions apart from plausible but hallucinated ones. The dataset includes two tasks: Spoken VQA (given an image and spoken question with options audio, choose the correct option) and Hallucination detection (given an image and three statements, decide for each statement whether it is True or False, with exactly one statement grounded). Each task is offered in two language tracks: English and Modern Standard Arabic (MSA), which are parallel with the same images and answers, and questions are translations of each other. The dataset spans 18 Arab countries, including Algeria, Bahrain, Egypt, etc. Audio in the train, dev, and devtest splits is synthetically generated using TTS, while the final test set uses human-recorded audio. Files include images, audio, and JSONL files with fields such as ID, image path, audio path, and labels. Splits consist of train (3000 items with labels), dev (500 items with labels), devtest (500 items without labels), and test (1000 items without labels). Evaluation metrics are separate for each track: accuracy for Spoken VQA and contrastive instability (CI) for Hallucination detection. Baseline models and example notebooks are provided, and the dataset is licensed under CC BY-NC-SA 4.0.




