OST-Bench
收藏资源简介:
OST-Bench 是一个用于评估在线时空场景理解能力的基准数据集,由上海人工智能实验室等机构创建。该数据集由 1.4k 个场景和 10k 个问答对组成,数据来源于 ScanNet、Matterport3D 和 ARKitScenes。数据集的设计模拟了真实世界中的场景,要求模型具备在线时空推理能力,能够根据动态更新的视觉信息进行推理。OST-Bench 的引入填补了现有基准数据集在在线时空推理能力评估方面的空白,为评估多模态大型语言模型在现实世界场景中的表现提供了重要的测试平台。
OST-Bench is a benchmark dataset for evaluating online spatio-temporal scene understanding capabilities, developed by institutions including the Shanghai AI Laboratory. It comprises 1.4k scenes and 10k question-answer pairs, sourced from ScanNet, Matterport3D and ARKitScenes. The dataset is designed to simulate real-world scenarios, requiring models to possess online spatio-temporal reasoning abilities and conduct reasoning based on dynamically updated visual information. The introduction of OST-Bench fills the gap in existing benchmark datasets for evaluating online spatio-temporal reasoning capabilities, providing a critical testbed for assessing the performance of multimodal large language models in real-world scenarios.
OST-Bench 数据集概述
基本信息
- 数据集名称: OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
- 作者: Jingli Lin*, Chenming Zhu*, Runsen Xu, Xiaohan Mao, Xihui Liu, Tai Wang†, Jiangmiao Pang†
- 贡献: *Equal Contribution †Corresponding Author
- 相关链接: Paper | Code | Data | arXiv
数据集简介
- 目的: 评估多模态大语言模型(MLLMs)在在线时空场景理解中的能力。
- 特点: 基于主动探索场景的视角,评估模型的在线时空理解能力。
- 数据来源: ScanNet, Matterport3D, ARKitScenes
- 规模: 1.4k场景和10k问答对
数据集分类
- 主要类别: 3个主要问题类别
- 子类别: 15个细粒度问题子类型
数据集统计
- 子类型分布: 包含详细统计
- 词云: 展示高频词汇
- 对话长度分布: 展示对话长度的统计信息
示例场景
- 场景选择: 1mp3d_0030_region2scene0050_00scene0100_00
- 系统提示: 假设你正在探索一个房间,所有物体都是静止的。随着时间的推移,你改变在房间中的位置和方向,并拍摄图像。
- 对话轮次: 5轮
- 问题示例: "Remember, have you seen any flowerpot(s) so far?"
- 选项: A. Yes, B. No
排行榜
- 评估模型: 包括专有模型和开源模型
- 评估指标: Overall, Agent State, Agent Visible Info, Agent Object Spatial
- 基准参考: Human-Level, Chance-Level
分析
- 探索过程中的性能下降: 模型准确性随着探索的深入显著下降。
- 错误分布统计: 推理错误占所有错误的60%以上。
- 时空推理捷径: 模型倾向于避免检索关键信息,依赖浅层推理。
- 跨视图分析: 模型在复杂线索或长期记忆检索需求下性能显著下降。
引用
bibtex @article{lin2025ostbench, title={OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding}, author={JingLi Lin and Chenming Zhu and Runsen Xu and Xiaohan Mao and Xihui Liu and Tai Wang and Jiangmiao Pang}, journal={arXiv preprint arXiv:2507.07984}, year={2025} }




