MiGA
收藏资源简介:
MiGA是一个大规模多夹爪感知数据集,由利物浦大学等机构联合创建,旨在填补当前VLA模型因忽略夹爪形态差异而导致的策略学习空白。该数据集包含103,000条轨迹,覆盖平行夹爪、真空吸盘、三指夹爪、软体夹爪和灵巧手五种类型,涉及36项任务,配备多视角RGB-D观测、本体状态和子步骤分解,并包含约5%的失败演示以分析夹爪能力边界。数据通过人工遥操作在仿真和真实环境中采集,任务设计刻意引发形态依赖策略,确保同一目标下不同夹爪展现出截然不同的操作轨迹。该数据集可用于训练和评估夹爪感知的视觉语言动作模型,旨在解决机器人操作中由夹爪形态差异导致的策略泛化难题,推动具身智能向更鲁棒和自适应的方向发展。
MiGA is a large-scale multi-gripper perception dataset co-developed by the University of Liverpool and other institutions, aimed at filling the gap in policy learning in current Vision-Language-Action (VLA) models caused by their neglect of gripper morphology differences. This dataset contains 103,000 trajectories covering five types of grippers: parallel jaw grippers, vacuum suction cups, 3-finger grippers, soft grippers, and dexterous hands, spanning 36 tasks. It is equipped with multi-view RGB-D observations, proprioceptive states, and sub-step decompositions, and includes approximately 5% of failure demonstrations to analyze the performance boundaries of different grippers. The data is collected via manual teleoperation in both simulation and real-world environments. The task design intentionally induces morphology-dependent policies, ensuring that different grippers exhibit distinctly different manipulation trajectories for the same target. This dataset can be used to train and evaluate vision-language-action models for gripper perception, aiming to address the policy generalization challenge in robotic manipulation caused by differences in gripper morphology, and advancing embodied intelligence towards more robust and adaptive directions.
GVLA: Gripper-aware Vision Language Action Models
数据集概述
GVLA 项目提出并发布了 MiGA (Multi-Gripper-Aware) 数据集,这是一个面向多夹爪感知的机器人操作数据集。该研究已被 ECCV 2026 接收。
MiGA 数据集核心信息
| 特性 | 详情 |
|---|---|
| 规模 | 103,000 条演示轨迹 |
| 夹爪类型 | 5 种不同夹爪类型 |
| 任务数量 | 36 个任务 |
| 环境 | 涵盖仿真环境与真实世界机器人设置 |
| 多机器人平台 | 数据来自多种机器人平台 |
数据集特点
- 明确捕捉在相同任务目标下,不同夹爪类型(如平行夹爪与吸盘)所需的不同接触选择、接近方向和执行策略
- 任务类别涵盖:分离(singulated)、堆叠(stacked)、受限(constrained)和语义(semantic)任务
- 在任务类别和夹爪覆盖范围上保持平衡
- 弥补了现有 VLA 数据集主要依赖平行夹爪、缺乏多夹爪感知学习的空白
GVLA 方法核心内容
GVLA 通过两种关键机制将夹爪本体信息注入 VLA 模型:
-
多夹爪分词器(Multi-gripper Tokenization):将机器人平台、夹爪类型和夹爪实例编码为可学习的 token,提供结构化条件信息。
-
双混合适配器(Dual Mixture-of-Adapters):通过平台感知和夹爪感知的适配器专家池,将计算路由到特定本体的策略。
实验结果摘要
- 在仿真和真实世界验证中,GVLA 均优于当前基线方法
- 提升了夹爪感知操作能力
- 改善了对新物体的零样本泛化能力
- 支持对未见任务或新夹爪的少样本适应
- 在 UR5 机器人搭配 Robotiq 2F-85 夹爪的真实世界验证中展示了更优的适应性能
相关资源
- 论文预印本(arXiv)
- 数据集
- 脚本代码
引用信息
该工作由利物浦大学、穆罕默德·本·扎耶德人工智能大学、印度科学研究所、东京大学、ZHAW、阿肯色大学、Physical Intelligence 和华中科技大学的研究人员合作完成,发表于 ECCV 2026。

- 1Gripper-aware Vision Language Action Models利物浦大学; 穆罕默德·本·扎耶德人工智能大学; 印度科学研究所; 东京大学; 苏黎世应用科技大学; 阿肯色大学; 物理智能公司; 华中科技大学 · 2026年



