BEAR
收藏资源简介:
BEAR数据集是一个全面的基准数据集,旨在评估多模态语言模型(MLLMs)的具身能力。该数据集由6个类别和14个原子技能组成,包含4469个交织的图像-视频-文本条目,涵盖了从低级指向、轨迹理解、空间推理到高级规划等任务。数据集的创建过程涉及从13个不同的数据源收集数据,并通过多阶段的生成流程和人工验证确保数据的多样性和准确性。BEAR数据集旨在帮助研究人员评估和改进MLLMs的具身能力,并推动具身智能领域的发展。
The BEAR dataset is a comprehensive benchmark dataset designed to evaluate the embodied capabilities of multimodal large language models (MLLMs). Comprising 6 categories and 14 atomic skills, the dataset contains 4,469 interleaved image-video-text entries, covering tasks ranging from low-level pointing, trajectory understanding, spatial reasoning to high-level planning. The dataset was created by collecting data from 13 distinct data sources, and its diversity and accuracy were ensured through a multi-stage generation pipeline and manual verification. The BEAR dataset aims to assist researchers in evaluating and improving the embodied capabilities of MLLMs, and promote the development of the embodied intelligence field.
BEAR 数据集概述
数据集基本信息
- 数据集名称: BEAR (Benchmarking and Enhancing Multimodal Language Models for Atomic Embodied Capabilities)
- 数据规模: 4,469个交错的图像-视频-文本VQA样本
- 类别数量: 6个主要类别
- 子类型数量: 15个细分子类型
- 问题类型分布:
- 多项选择题: 57.4%
- 自由形式问题: 42.6%
- 新生成样本: 93.3%
评估能力类别
基础能力类别
- Pointing (指向)
- Bounding Box (边界框)
- Trajectory Reasoning (轨迹推理)
- Spatial Reasoning (空间推理)
- Task Planning (任务规划)
长视野类别
- 从AI2-THOR模拟器中收集的35个情景
- 将具身情景分解为技能导向的步骤进行离线评估
- 涵盖规划、物体搜索、导航、空间推理、感知和放置等步骤
模型评估结果
整体性能对比
- 专有模型平均分: 39.2
- 开源模型平均分: 25.8
- 性能差距: 13.4
评估模型数量
- 总模型数: 20个代表性MLLMs
- 开源模型: 12个
- 专有模型: 8个
性能指标缩写说明
- GEN: General Object (Pointing/Box)
- SPA: Spatial Object (Pointing/Box)
- PRT: Semantic Part (Pointing/Box)
- PRG: Task Process Reasoning
- PRD: Next Action Prediction
- GPR: Gripper Trajectory Reasoning
- HND: Human Hand Trajectory Reasoning
- OBJ: Object Trajectory Reasoning
- LOC: Object Localization
- PTH: Path Planning
- DIR: Relative Direction
BEAR-Agent 增强方案
- 类型: 多模态可对话智能体
- 功能: 利用视觉工具增强MLLMs的具身能力
- 效果: 显著提升InternVL3-14B和GPT-5在BEAR基准上的性能

- 1BEAR: Benchmarking and Enhancing Multimodal Language Models for Atomic Embodied CapabilitiesNortheastern University, The Chinese University of Hong Kong, Peking University, Westlake University, Harvard University, Purdue University, University of Oxford · 2025年



