ReasonVQA
收藏资源简介:
ReasonVQA是一个针对视觉问答任务的新数据集,它集成了结构化的百科全书知识,并通过低成本框架构建,能够生成复杂的多跳问题。该数据集包含大量问题,分为1跳、2跳和3跳三个复杂度级别,要求模型具备强大的多跳推理能力。数据集构建过程包括外部知识整合、问题生成和数据集构建三个步骤。数据集利用了Wikidata和Visual Genome等知识库和图像数据源,并通过模板生成问题和选项,同时进行了答案分布平衡和数据集分割,以减少偏差并提高模型的挑战性。
ReasonVQA is a novel dataset tailored for visual question answering (VQA) tasks. It integrates structured encyclopedic knowledge and is built using a low-cost framework, enabling the generation of complex multi-hop questions. This dataset contains a large volume of questions categorized into three complexity levels: 1-hop, 2-hop, and 3-hop, which require models to possess strong multi-hop reasoning capabilities. The development pipeline of this dataset includes three key steps: external knowledge integration, question generation, and dataset construction. The dataset leverages knowledge bases and image data sources such as Wikidata and Visual Genome, generates questions and answer options via template-based methods, and conducts answer distribution balancing and dataset splitting to reduce biases and enhance the challenge for models.

- 1ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering博世人工智能中心, 柏林工业大学, 弗劳恩霍夫FOKUS研究所 · 2025年



