DriveBench
收藏资源简介:
DriveBench是由上海人工智能实验室等机构创建的一个基准数据集,旨在评估视觉语言模型(VLMs)在自动驾驶任务中的可靠性。该数据集包含19,200帧图像和20,498个问答对,涵盖感知、预测、规划和解释等四大主流驾驶任务,并在17种不同的设置(如干净、损坏和纯文本输入)下进行评估。数据集的内容包括多种问题类型(如多选题、开放式问题和视觉基础问题),数据来源广泛,涵盖了真实世界的自动驾驶场景。数据集的创建过程包括对现有驾驶数据集的深入分析,并通过重新采样解决了数据分布不平衡的问题。DriveBench的应用领域主要集中在自动驾驶领域,旨在揭示VLMs在视觉基础和多模态推理方面的局限性,并推动更可靠、可解释的自动驾驶决策系统的发展。
DriveBench is a benchmark dataset developed by institutions including the Shanghai AI Laboratory, designed to evaluate the reliability of Vision-Language Models (VLMs) in autonomous driving tasks. This dataset contains 19,200 image frames and 20,498 question-answer pairs, covering four mainstream driving tasks including perception, prediction, planning, and explanation, and is evaluated across 17 distinct settings such as clean, corrupted, and pure text input conditions. Its content encompasses diverse question types, including multiple-choice questions, open-ended questions, and visual grounding questions, with extensive data sources covering real-world autonomous driving scenarios. The construction of DriveBench involves in-depth analysis of existing driving datasets, and resolves the issue of imbalanced data distribution through resampling. DriveBench's application scenarios are primarily centered on the autonomous driving domain, aiming to uncover the limitations of VLMs in visual grounding and multimodal reasoning, and advance the development of more reliable and interpretable autonomous driving decision-making systems.




