QLEVR
收藏资源简介:
合成数据集已成功用于探索视觉问答数据集的推理能力。例如,CLEVR 测试一系列视觉推理能力。 CLEVR 中的问题侧重于形状、颜色和大小的比较、数字推理和存在主张。本文介绍了一个最小偏差、诊断性视觉问答数据集 QLEVR,它超越了存在和数值量化,专注于更复杂的量词及其组合,例如,询问是否有两个以上的红球小于至少图像中的三个蓝色球。我们描述了数据集是如何创建的,并对最先进的视觉问答模型进行了首次评估,表明 QLEVR 对我们当前的模型提出了巨大的挑战。描述和图片来自:QLEVR 数据集生成
Synthetic datasets have been successfully used to explore the reasoning capabilities of visual question answering (VQA) datasets. For example, CLEVR tests a range of visual reasoning capabilities. Questions in CLEVR focus on comparisons of shape, color, and size, numerical reasoning, and existential claims. This paper introduces QLEVR, a minimally biased, diagnostic visual question answering dataset that extends beyond existential and numerical quantification, focusing on more complex quantifiers and their combinations. For instance, it queries whether there are more than two red balls smaller than at least three blue balls in the image. We describe the creation process of the dataset and conduct the first evaluation of state-of-the-art visual question answering models, demonstrating that QLEVR poses substantial challenges to our current models. Descriptions and images are sourced from: QLEVR Dataset Generation




