GEM-4M
收藏资源简介:
GEM-4M是由清华大学与腾讯混元团队联合构建的大规模高质量具身智能预训练数据集,旨在为生成式监督的视觉语言模型提供深度感知与物理推理支持。该数据集包含约400万条问答对,数据来源融合了具身任务中的基础定位、时空规划与物理推理等多模态信息,并配以高质量的深度监督信号。其构建过程通过精心设计的数据引擎整合了多样化的具身任务数据,以强化模型对场景几何结构与物理约束的理解。该数据集主要应用于提升具身视觉语言模型在真实物理环境中的语义理解与操作能力,旨在解决传统模型在高级语义推理与低级空间物理知识之间的脱节问题,推动机器人自主任务执行的发展。
GEM-4M is a large-scale high-quality embodied intelligence pre-training dataset jointly constructed by Tsinghua University and Tencent Hunyuan Team. It is designed to provide deep perception and physical reasoning support for generative-supervised visual-language models. This dataset contains approximately 4 million question-answer pairs, with data sources integrating multimodal information including basic localization, spatio-temporal planning and physical reasoning from embodied tasks, and is paired with high-quality deep supervision signals. During the construction process, diverse embodied task data are integrated through a meticulously designed data engine, so as to enhance the model's understanding of scene geometric structures and physical constraints. This dataset is primarily applied to improve the semantic understanding and manipulation capabilities of embodied visual-language models in real physical environments. It aims to address the gap between advanced semantic reasoning and low-level spatial physical knowledge in traditional models, and promote the development of autonomous task execution for robots.
数据集概述:GEM-4M
数据集名称:GEM-4M(Generative-supervised Embodied vision-language Model - 4 Million)
发布机构:清华大学、腾讯混元
核心定位:一个大规模、高质量的数据集,专为生成式监督的具身视觉语言模型(VLM)设计,旨在弥合高层语义理解与低层空间物理知识之间的鸿沟。
数据规模与组成:
- 规模:包含约400万(4M)数据样本。
- 内容类型:混合了基础(Grounding)、推理(Reasoning) 与规划(Planning) 三类数据。
- 监督信号:每份数据均配套高质量深度图(Depth supervision)。
数据用途: 该数据集用于在VLM预训练阶段引入深度图生成任务作为辅助目标。通过联合训练语言建模与深度生成目标,增强模型在具身智能任务中的物理 grounding 能力和语义推理能力。基于GEM-4M训练的GEM-VLA模型在LIBERO等基准测试和真实机器人操作任务中取得了领先表现。

- 1GEM: Generative Supervision Helps Embodied Intelligence清华大学; 腾讯·混元 · 2026年



