ReasonCore/open-spatial-reasoning
收藏资源简介:
Open Spatial Reasoning是一个用于评估从单张驾驶图像中进行3D空间推理的多选题数据集。该数据集旨在揭示前沿视觉语言模型在复杂驾驶场景(如高架道路、斜坡、弯道和交叉口)中依赖像素布局启发式方法(例如图像中位置较低表示更近、边界框更大表示更近)进行错误推理的失败模式。数据集包含由PlusAI运营的自动驾驶车辆收集的驾驶场景图像,每张图像带有编号的边界框,问题涉及度量3D推理,包括距离、横向位置、排序和方向。每个样本包括一张图像、一个问题、四个答案选项和正确答案。问题类别涵盖识别绝对距离、相对距离、选择更近物体、识别最右侧物体、左到右排序、识别位置和识别方向等任务类型。
Open Spatial Reasoning is a multiple-choice dataset for evaluating 3D spatial reasoning from single driving images. It is designed to surface the failure mode of frontier vision-language models that often answer questions correctly by luck while reasoning incorrectly, leaning on pixel-layout heuristics (e.g., lower in the frame = closer, bigger box = nearer) that break down on elevated roads, slopes, curves, and intersections. The dataset includes driving-scene images collected by autonomous vehicles operated by PlusAI, each with numbered bounding boxes referencing objects in the scene. Questions probe metric 3D reasoning about distance, lateral position, ordering, and heading. Each sample pairs an image with a question, four answer choices, and the correct answer letter. Question categories include identifying absolute distance, relative distance, picking closer objects, identifying the rightmost object, ordering leftmost objects, identifying position, and identifying heading.




