InsScene4D-147K
收藏资源简介:
InsScene4D-147K是由地平线机器人与南洋理工大学等机构联合构建的大规模4D场景理解数据集,旨在支持几何与实例联合学习的训练与评估。该数据集涵盖真实与合成、静态与动态场景,包含147,000个数据样本,总计约6,000万个实例掩码,数据来源包括RGB图像、深度图、相机位姿及点云,并通过自动化几何引导标注流程生成时序一致的实例标签。其创建过程采用基于空间基础模型的几何重建、投影与掩码细化流水线,实现了多视角一致的几何与实例标注的高效生成。该数据集主要应用于在线4D场景理解、实例空间跟踪、开放词汇分割等下游任务,旨在解决动态环境中长期序列的几何-实例一致性预测缺乏高质量监督数据的问题。
InsScene4D-147K is a large-scale 4D scene understanding dataset jointly constructed by Horizon Robotics, Nanyang Technological University and other institutions, aiming to support the training and evaluation of joint learning of geometry and instance. This dataset covers real and synthetic, static and dynamic scenes, contains 147,000 data samples totaling approximately 60 million instance masks. Its data sources include RGB images, depth maps, camera poses and point clouds, and temporally consistent instance labels are generated through an automated geometry-guided annotation pipeline. Its creation process adopts a spatial foundation model-based pipeline for geometry reconstruction, projection and mask refinement, enabling efficient generation of multi-view consistent geometry and instance annotations. This dataset is mainly applied to downstream tasks such as online 4D scene understanding, instance spatial tracking and open-vocabulary segmentation, and aims to address the problem of lack of high-quality supervised data for geometry-instance consistency prediction in long-term sequences in dynamic environments.
数据集概述
数据集名称
InsScene4D-147K
数据规模
总计 147,000 条序列,涵盖真实与合成场景、静态与动态场景。
数据构成
- RGB 图像:包含对应的深度图、相机位姿(pose)。
- 实例掩码(Instance Masks):通过自动化的几何引导标注流程生成,具有时间一致性(temporally consistent)。
数据来源
数据集由三个互补的数据源构成:
- RoboTwin 2.0:利用官方流程生成带有时间一致性实例掩码的 RGB-D 序列。
- RealEstate10K:使用 DA3 深度和真实位姿,通过 TSDF 网格重建进行 ID 继承,并结合 SAM2 自动及框提示掩码进行标注优化。
- HOI4D:在静态网格重建前移除动态区域,并利用提供的动态对象注释更新投影伪标签。
关键特点
- 提供 4D 一致的实例注释。
- 旨在解决高质量 4D 监督数据缺乏的问题。
- 为 IGGT4D 模型(一种流式实例几何Transformer)的在线 4D 场景理解提供训练和评估基准。
应用场景
- 3D 重建
- 位姿估计
- 实例空间追踪
- 开放词汇语义分割
- 4D 场景问答




