SIBench
收藏资源简介:
SIBench是一个包含近20个开源数据集的视觉空间推理评估基准,涵盖了23种不同的视觉空间推理任务设置。这些任务设置旨在评估视觉语言模型(VLMs)在空间推理方面的能力,包括基本感知、空间理解和规划三个层次的能力。SIBench提供了对现有模型在视觉空间推理任务中的表现进行综合评估的工具,揭示了当前模型在精确数值估计、多视图推理、时间信息处理和空间想象等方面的不足。
SIBench is a visual spatial reasoning evaluation benchmark comprising nearly 20 open-source datasets and covering 23 distinct visual spatial reasoning task settings. These task settings are designed to evaluate the spatial reasoning capabilities of Vision-Language Models (VLMs), including three hierarchical ability dimensions: basic perception, spatial understanding, and planning. SIBench provides a toolkit for comprehensively assessing the performance of existing models on visual spatial reasoning tasks, and reveals the shortcomings of current models in precise numerical estimation, multi-view reasoning, temporal information processing, and spatial imagination.




