DRIVINGVQA
收藏资源简介:
DRIVINGVQA是由瑞士洛桑联邦理工学院创建的一个视觉问答数据集,旨在评估和改进视觉语言模型在复杂现实场景中的视觉链式推理能力。该数据集包含3931个样本,每个样本包括一个或多个多项选择题、相关实体的边界框坐标以及与视觉内容对齐的解释。数据来源于驾驶理论考试,涵盖了广泛的驾驶场景。数据集的创建过程包括数据收集、相关实体的人工标注以及生成与视觉元素相关的解释。DRIVINGVQA的应用领域主要集中在自动驾驶和视觉推理任务中,旨在解决多对象识别、空间关系推理和决策制定等问题。
DRIVINGVQA is a visual question answering (VQA) dataset developed by École Polytechnique Fédérale de Lausanne (EPFL), which aims to evaluate and enhance the visual chain-of-reasoning capabilities of vision-language models in complex real-world scenarios. This dataset comprises 3931 samples, each containing one or more multiple-choice questions, bounding box coordinates of relevant entities, and explanations aligned with the visual content. The data is sourced from driving theory examinations and covers a broad spectrum of driving scenarios. The dataset creation workflow includes data collection, manual annotation of relevant entities, and generation of explanations linked to visual elements. The application domains of DRIVINGVQA primarily focus on autonomous driving and visual reasoning tasks, with the goal of addressing challenges such as multi-object recognition, spatial relation reasoning and decision-making.

- 1DRIVINGVQA: Analyzing Visual Chain-of-Thought Reasoning of Vision Language Models in Real-World Scenarios with Driving Theory Tests瑞士洛桑联邦理工学院 · 2025年



