GroundVQA
收藏资源简介:
GroundVQA是一个对抗性视觉问答(VQA)基准数据集,用于评估前沿多模态大语言模型(MLLMs)在视觉基础、抗幻觉、不确定性校准、OCR鲁棒性和物理基础推理等方面的能力。该数据集包含约4,000个经过手动审核的对抗性VQA示例,这些示例基于作者图像以及来自Visual Genome、WearVQA、TextCaps、DocVQA、ChartQA和GQA等多个公共图像源的数据构建。与传统的VQA基准不同,GroundVQA侧重于模型评估而非训练,要求模型能够拒绝不支持的前提、在视觉证据不足时主动弃权、区分可见观察内容与推断内容,并避免对对象、属性、关系或OCR文本产生幻觉。数据集涵盖多个推理类别,包括OCR基础、空间推理、物理推理、对抗性幻觉、文档理解、图表推理以及模糊视觉证据处理。GroundVQA适用于前沿MLLMs的评估、幻觉分析、视觉基础研究以及不确定性感知多模态推理的基准测试,但不适用于监督训练任务。
GroundVQA is an adversarial Visual Question Answering (VQA) benchmark dataset developed to evaluate the capabilities of state-of-the-art multimodal large language models (MLLMs) across visual grounding, hallucination mitigation, uncertainty calibration, OCR robustness, and physical grounding reasoning. This dataset contains approximately 4,000 manually reviewed adversarial VQA examples, constructed using both author-provided images and data from multiple public image sources including Visual Genome, WearVQA, TextCaps, DocVQA, ChartQA, and GQA. Unlike traditional VQA benchmarks, GroundVQA focuses on model evaluation rather than training, requiring models to reject unsupported premises, actively abstain when visual evidence is insufficient, distinguish between observable content and inferred content, and avoid hallucinating objects, attributes, relationships, or OCR text. The dataset covers multiple reasoning categories, including OCR grounding, spatial reasoning, physical reasoning, adversarial hallucination, document understanding, chart reasoning, and ambiguous visual evidence handling. GroundVQA is suitable for evaluating state-of-the-art MLLMs, conducting hallucination analysis, advancing visual grounding research, and benchmarking uncertainty-aware multimodal reasoning, but it is not intended for supervised training tasks.
GroundVQA 数据集概述
数据集名称:GroundVQA
许可证:Apache-2.0
任务类别:视觉问答(Visual Question Answering)
语言:英语
规模:约 1,000 至 10,000 条样本(当前版本约 4,000 条)
数据集描述
GroundVQA 是一个对抗性视觉问答(VQA)基准,专为评估前沿多模态大语言模型(MLLMs)在以下方面的能力而设计:
- 视觉定位:模型对图像中具体元素的定位准确性。
- 抗幻觉能力:拒绝无依据的假设、在视觉证据不足时拒绝回答、区分可见观察与推断内容。
- 不确定性校准:在模棱两可的视觉证据下做出合理判断。
- OCR 鲁棒性:处理文本光学字符识别任务的稳健性。
- 物理推理:基于物理常识的推理能力。
数据集由约 4,000 条人工审核的对抗性 VQA 样本构成,图像来源包括作者自有图像以及 Visual Genome、WearVQA、TextCaps、DocVQA、ChartQA 和 GQA 等公开数据集(其中部分问答内容已重新整理)。
推理类别
GroundVQA 涵盖多种推理类型:
- OCR 定位:识别并定位图像中的文本。
- 空间推理:理解物体间的空间关系。
- 物理推理:基于物理规则的推断。
- 对抗性幻觉:检测模型对虚构信息的倾向。
- 文档理解:处理文档图像中的信息。
- 图表推理:分析图表数据与视觉呈现。
- 模糊视觉证据:在信息不明确时进行判断。
预期用途
- 评估前沿多模态大语言模型:作为测试基准。
- 幻觉分析:研究模型生成虚假信息的现象。
- 视觉定位研究:探索模型对图像元素的精确指向能力。
- 不确定性感知推理基准测试:推动模型在不确定性场景下的推理研究。
注意:该数据集主要用于评估,不适合作为监督训练数据。




