EmbSpatial-Bench
收藏资源简介:
EmbSpatial-Bench是由复旦大学数据科学与计算机科学学院创建的一个用于评估大型视觉语言模型(LVLMs)在具身任务中空间理解能力的基准数据集。该数据集包含3640个多选题,覆盖294个对象类别和6种空间关系,数据来源于MP3D、ScanNet和AI2-THOR等具身3D场景。创建过程涉及从3D场景中自动提取空间关系并生成问答对。EmbSpatial-Bench旨在解决LVLMs在具身环境中空间理解能力的评估问题,为具身AI系统的发展提供关键支持。
EmbSpatial-Bench is a benchmark dataset created by the School of Data Science and Computer Science, Fudan University, for evaluating the spatial comprehension capabilities of large vision-language models (LVLMs) in embodied tasks. This dataset contains 3,640 multiple-choice questions, covering 294 object categories and 6 types of spatial relationships, with data sourced from embodied 3D scenes such as MP3D, ScanNet, and AI2-THOR. Its creation process involves automatically extracting spatial relationships from 3D scenes and generating question-answer pairs. EmbSpatial-Bench aims to address the problem of evaluating the spatial comprehension abilities of LVLMs in embodied environments, providing critical support for the development of embodied AI systems.

- 1EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models复旦大学数据科学与计算机科学学院 · 2024年



