zzzrw/GEM-250K
收藏资源简介:
GEM-250K(或相关GEM-4M)是一个大规模多模态数据集,旨在支持生成式监督的具身视觉语言模型(GEM)的预训练。该数据集通过集成深度图生成任务,弥合了高级语义理解与低级空间物理知识之间的差距,以增强具身智能。数据集包含三种主要类型的数据:具身 grounding(涉及图像、问题和答案,用于基础感知)、空间推理(涉及视频、问题和答案,用于空间理解)以及时空规划(涉及图像、问题和答案,用于规划任务)。所有数据均配有高质量深度监督信息,帮助模型学习物理环境中的空间关系。数据集结构包括多个分割(spatiotemporal_planning, embodied_grounding, spatial_reasoning),总计超过25万个示例,总数据量约2.5 GB,支持图像-文本到文本的任务分类,语言为英语。
GEM-250K (or related GEM-4M) is a large-scale multimodal dataset designed to support the pre-training of Generative-supervised Embodied vision-language Models (GEM). By integrating a depth map generation task, the dataset bridges the gap between high-level semantic understanding and low-level spatial and physical knowledge, enhancing embodied intelligence. It comprises three main types of data: embodied grounding (involving images, questions, and answers for basic perception), spatial reasoning (involving videos, questions, and answers for spatial understanding), and spatiotemporal planning (involving images, questions, and answers for planning tasks). All data are paired with high-quality depth supervision to aid models in learning spatial relationships in physical environments. The dataset structure includes multiple splits (spatiotemporal_planning, embodied_grounding, spatial_reasoning), totaling over 250,000 examples with an overall size of approximately 2.5 GB, supporting image-text-to-text task categories, and the language is English.





