SpatialBench
收藏资源简介:
SpatialBench是一个基准测试套件,旨在评估多模态大语言模型在视频空间理解方面的能力。该数据集涵盖5个主要类别和15个子类别的空间任务,包括观察与测量、拓扑与组合、符号视觉推理、空间因果关系和空间规划。
SpatialBench is a benchmark suite designed to evaluate the video-based spatial understanding capabilities of multimodal large language models (LLMs). This dataset covers spatial tasks across 5 major categories and 15 subcategories, including observation and measurement, topology and composition, symbolic visual reasoning, spatial causality, and spatial planning.
SpatialBench 数据集概述
数据集基本信息
- 名称:SpatialBench
- 类型:视频空间理解基准测试数据集
- 用途:评估多模态大语言模型在视频空间认知方面的能力
- 关联论文:CVPR 2026论文《SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition》
数据集特征
- 多维度评估:涵盖5大类别和15个子类别的空间任务
- 观察与测量
- 拓扑与组合
- 符号视觉推理
- 空间因果关系
- 空间规划
数据集文件组成
- QA.txt:标准基准数据集,包含空间推理问题
- QA_fewshot.txt:专为"深度引导"模式设计的数据集变体
- test_sample.txt:用于快速测试和调试的小样本数据集
- dataset/:测试视频文件目录
数据格式
- 输入格式:JSON格式,包含样本对象列表
- 样本字段:
- problem_id:问题ID
- path:视频文件路径
- problem_type:问题类型
- problem:问题描述
- options:选项列表
- solution:标准答案
- scene_type:场景类型
评估方法
- 多选题:匹配模型输出选项,正确得1分,错误得0分
- 回归问题:使用平均相对准确率算法,得分范围0-1
- 加权总分:根据不同任务类别的难度和重要性进行加权计算
获取方式
- 下载地址:https://huggingface.co/datasets/XPR2004/SpatialBench
- 下载要求:需要安装Git LFS来下载视频文件
引用信息
bibtex @misc{xu2025spatialbenchbenchmarkingmultimodallarge, title={SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition}, author={Peiran Xu and Sudong Wang and Yao Zhu and Jianing Li and Yunjian Zhang}, year={2025}, eprint={2511.21471}, archivePrefix={arXiv}, primaryClass={cs.AI}, url={https://arxiv.org/abs/2511.21471}, }




