UIUC D3Field
收藏资源简介:
D3Fields 是由伊利诺伊大学厄巴纳-香槟分校和斯坦福大学联合开发的一种新型动态三维描述场(Dynamic 3D Descriptor Fields)数据集,旨在为机器人操作任务提供零样本(zero-shot)泛化能力。该数据集通过多视角 RGB-D 图像输入,利用基础模型(如 Grounding-DINO、SAM 等)提取语义特征和实例掩码,并将这些信息映射到三维空间中的任意点,形成动态、语义化的三维描述场。数据集内容丰富,涵盖多种日常操作任务(如整理鞋子、收集碎片、整理办公桌等)的场景表示和目标图像,数据来源为真实世界和模拟环境中的机器人操作实验。创建过程通过多视角融合和插值技术,无需额外训练即可生成描述场。该数据集的应用领域是机器人操作任务的零样本泛化,旨在解决传统方法在新任务和场景下的泛化难题,支持机器人根据二维目标图像完成多样化、动态化的三维操作任务。
D3Fields is a novel Dynamic 3D Descriptor Fields dataset co-developed by the University of Illinois Urbana-Champaign and Stanford University, which aims to provide zero-shot generalization capabilities for robotic manipulation tasks. This dataset takes multi-view RGB-D images as input, utilizes foundation models such as Grounding-DINO and SAM to extract semantic features and instance masks, and maps this information to arbitrary points in 3D space to form dynamic, semantic 3D descriptor fields. The dataset is rich in content, covering scene representations and target images for various daily manipulation tasks including sorting shoes, collecting debris, tidying desks and more. Its data is sourced from real-world and simulated robotic manipulation experiments. The creation process adopts multi-view fusion and interpolation techniques to generate descriptor fields without requiring additional training. The application scenario of this dataset focuses on zero-shot generalization for robotic manipulation tasks, aiming to address the generalization challenges faced by traditional methods in novel tasks and scenarios, and enables robots to complete diverse and dynamic 3D manipulation tasks based on 2D target images.




