Deform360
收藏资源简介:
Deform360是由布朗大学、哥伦比亚大学等机构联合创建的大规模多视角视觉触觉数据集,旨在为可变形物体世界建模提供高质量的基准数据。该数据集囊括了198种日常物体,包含1,980个交互序列,总计超过2,150万帧图像和215.7小时的多视角视频,并同步采集了双工触觉信号,数据规模极为庞大。其构建过程通过41个环绕摄像头和配备触觉的UMI夹爪系统,捕获了物体全局运动和接触引起的局部形变,并利用无标记三维跟踪流程生成了高保真的几何与运动标注。该数据集主要应用于机器人操控领域,旨在系统评估和比较二维视频与三维粒子世界模型在可变形物体动力学预测、接触检测及机器人规划任务中的性能,以解决当前因缺乏多样、大规模真实世界数据而难以深入理解模型优劣的核心挑战。
Deform360 is a large-scale multi-view visual-tactile dataset jointly created by Brown University, Columbia University and other institutions, aiming to provide high-quality benchmark data for modeling deformable object worlds. This dataset covers 198 daily objects, contains 1,980 interaction sequences, with a total of more than 21.5 million image frames and 215.7 hours of multi-view video, and simultaneously collected duplex tactile signals, resulting in an extremely large data scale. During its construction, 41 surround-view cameras and a tactile-equipped UMI gripper system were used to capture the global motion of objects and local deformations caused by contacts, and a markerless 3D tracking pipeline was utilized to generate high-fidelity geometric and motion annotations. This dataset is mainly applied in the field of robotic manipulation, aiming to systematically evaluate and compare the performance of 2D video and 3D particle world models in tasks including deformable object dynamics prediction, contact detection and robotic planning, so as to address the core challenge that it is currently difficult to deeply assess the strengths and weaknesses of models due to the lack of diverse, large-scale real-world data.
数据集概述
Deform360 是一个大规模多视角触觉数据集,专注于可变形物体,用于支持2D和3D世界模型的基准测试。
核心数据规模
- 物体数量: 198个日常可变形物体,涵盖17个类别。
- 交互序列: 1,980个交互序列。
- 总时长: 累计215.7小时的多视角视频。
- 总帧数: 约2,330万帧。
- 原始视频: 74,850个。
- 录制规格: 720p分辨率,30 FPS,同步捕捉。
传感器配置
- 视觉: 41个环绕相机,提供360°全覆盖。
- 触觉: 双手触觉夹爪(UMI平台)。
物体分类(按材料响应)
- 1D 线性物体: 绳索、电缆、电线,具有不同的刚度和厚度。
- 2D 薄壳物体: 布料、织物、纸张,包括多种纺织品、气囊和薄壳。
- 3D 体积物体: 毛绒玩具、泡沫、填充物,表现出较大的形状变化。
关键特性
- 标注: 通过无标记物视觉触觉跟踪流程,生成密集的3D粒子标注。
- 几何重建: 使用逐帧3D高斯泼溅(3D Gaussian Splatting)恢复高保真几何形状。
- 跟踪: 利用CoTracker3进行2D跟踪,并通过渲染深度反投影并融合41个视角的3D点。
- 优化: 结合物理信息的细化过程,包含形状、刚性、平滑度和触觉损失。
- 触觉信号: 同步的触觉传感器提供法向压力接触提示,在视觉遮挡时约束粒子运动。
基准测试结论
- 未来预测(已见物体,未见交互片段): ParticleFormer 在所有报告的指标上表现最佳。
- 未见物体上的视觉指标: 预训练的Cosmos模型在重建方面视觉优势最强,但可能偏离指令动作。
开源与许可
- 许可证: 数据集和完整流程(捕捉、重建、感知和世界模型基线)均以MIT许可证发布。
- 数据获取: 通过Hugging Face平台(
huggingface.co/datasets/brownu/deform360)访问。
作者与引用
- 作者: Hongyu Li, Wanjia Fu, Xiaoyan Cong, Zekun Li, Binghao Huang, Hanxiao Jiang, Xintong He, Yiqing Liang, Rao Fu, Tao Lu, Srinath Sridhar, Kevin A. Smith, George Konidaris, Yunzhu Li。
- 机构: 布朗大学、哥伦比亚大学、麻省理工学院。
- 引用: 已收录于ECCV 2026。

- 1Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models布朗大学; 哥伦比亚大学; 麻省理工学院 · 2026年



