PAC Bench
收藏资源简介:
PAC Bench是一个全面的数据集,旨在评估视觉-语言模型(VLMs)在执行操纵策略时的基本属性、可利用性和约束条件(PAC)的理解。数据集包含超过30,000个注释,包括673张真实世界图像(115个对象类别、15种属性类型、每个类别1-3个定义的可利用性),100个真实世界的拟人视角场景和120个独特的模拟约束场景,跨越四个任务。PAC Bench的数据采集和整理过程采用了多方面的方法,结合了来自现实世界和模拟的图像数据,确保了视觉的多样性和真实性。该数据集的创建是为了填补现有基准在评估VLMs对执行操纵行动的基本前提的理解方面的空白,并为构建更可靠和物理地面的机器人操纵模型提供指导。
PAC Bench is a comprehensive dataset designed to evaluate the understanding of Vision-Language Models (VLMs) regarding their grasp of fundamental properties, affordances, and constraints (PAC) when executing manipulation strategies. The dataset contains over 30,000 annotations, including 673 real-world images (115 object categories, 15 attribute types, and 1–3 defined affordances per category), 100 real-world egocentric scenes, and 120 unique simulated constraint scenarios, spanning four tasks. The data collection and curation workflow of PAC Bench adopts a multi-faceted approach combining real-world and simulated image data to ensure visual diversity and authenticity. This dataset was created to fill the gap in existing benchmarks for evaluating VLMs' understanding of the fundamental prerequisites for executing manipulation actions, and to offer guidance for building more reliable and physically grounded robotic manipulation models.
PAC Bench数据集概述
数据集简介
- 名称: PAC Bench
- 目的: 评估视觉语言模型(VLMs)在机器人操作任务中对物理属性(P)、功能可供性(A)和约束条件(C)的理解能力
- 核心评估维度:
- Properties: 物体固有特性(材料、重量等)
- Affordances: 动作可能性(可抓取、可堆叠等)
- Constraints: 物理限制(稳定性、可达性等)
数据集构成
- 总标注量: 超过30,000个
- 数据组成:
- 673张真实世界图像(115个物体类别)
- 100个人形机器人视角场景
- 120个模拟约束场景
- 属性类型: 15种
- 功能可供性: 每个类别定义1-3个
数据子集
-
Constraint Images Dataset
- 模拟场景测试物理约束理解
- 包含:
- 不可能放置
- 支撑/遮挡问题
- 可达性问题
- 稳定性约束
-
Humanoid Robot Dataset
- 从Unitree G1人形机器人视角采集的真实场景
-
Open Images Dataset
- 多样化真实图像用于属性和可供性评估
- 覆盖115个物体类别
-
RoboCasa Objects Dataset
- 家庭物品多角度视图(每个物体24个视角)
- 示例物体:
- 奶酪块
- 甜甜圈
- 法棍面包
评估结果
属性理解准确率(%)
| 模型 | Open Images | Humanoid | 平均 |
|---|---|---|---|
| Claude 3.5 Sonnet | 27.8 | 50.2 | 27.8 |
| Gemini 2.0 Flash 001 | 44.1 | 55.2 | 44.1 |
| GPT-4.1 | 42.4 | 51.2 | 42.4 |
| Llama 4 Maverick | 49.4 | 43.8 | 49.4 |
约束理解准确率(%)
| 模型 | 模拟 | 真实世界 | 平均 |
|---|---|---|---|
| Gemini 2.5 Pro P | 25.8 | 11.3 | 25.8 |
| GPT-4.1 | 13.6 | 11.3 | 13.6 |
| GPT-4.1 Mini | 4.4 | 18.8 | 4.4 |
| Llama 3.2 11B Vision I | 17.5 | 0.0 | 17.5 |
功能可供性识别
- 所有模型在识别全部正确可供性时表现接近零
- 例外:
- GPT-4.1在家居固定装置上达到20%
- Qwen 2.5 VL在工具和硬件上达到11.1%




