4D-VLA
收藏资源简介:
4D-VLA数据集由复旦大学数据科学学院和华为诺亚方舟实验室创建,旨在解决现有预训练模型中输入信息不完整导致的问题,例如坐标系统混乱和状态混乱。该数据集包含88条记录,旨在帮助机器人更好地理解空间和时间信息,以实现更准确和通用的行为控制。数据集的创建过程涉及将深度和时序信息整合到视觉特征中,以生成对齐机器人坐标系统和场景的4D时空表示。4D-VLA数据集在模拟和真实世界环境中进行了验证,并展示了其在机器人控制任务中的优越性能。
The 4D-VLA dataset was created by the School of Data Science at Fudan University and Huawei Noah's Ark Lab, aiming to resolve issues caused by incomplete input information in existing pre-trained models, such as coordinate system confusion and state confusion. This dataset includes 88 records, designed to help robots better understand spatial and temporal information to achieve more accurate and generalizable behavior control. The development of the 4D-VLA dataset involves integrating depth and temporal information into visual features to generate a 4D spatiotemporal representation aligned with the robot's coordinate system and the scene. The 4D-VLA dataset has been validated in both simulated and real-world environments, and has demonstrated superior performance in robot control tasks.
4D-VLA数据集概述
基本信息
- 数据集名称: 4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration
- 相关论文: arXiv:2506.22242
- 作者:
- Jiahui Zhang<sup>1*</sup>
- Yurui Chen<sup>1*</sup>
- Yueming Xu<sup>1</sup>
- Ze Huang<sup>1</sup>
- Yanpeng Zhou<sup>2</sup>
- Yu-Jie Yuan<sup>2</sup>
- Xinyue Cai<sup>2</sup>
- Guowei Huang<sup>2</sup>
- Xingyue Quan<sup>2</sup>
- Hang Xu<sup>2</sup>
- Li Zhang<sup>1</sup>
- 机构:
- <sup>1</sup>Fudan University
- <sup>2</sup>Huawei Noah’s Ark Lab
数据集特点
- 设计哲学: 强调先前方法在输入中缺乏准确动作推断的关键线索,导致目标动作分布具有高方差或非平滑性。
- 验证环境: 在模拟和真实机器人环境中验证方法性能。
- 基准对比: 包含OpenVLA基线和4D-VLA方法的性能报告。
引用信息
bibtex @article{zhang2025vla, title={4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration}, author={Zhang, Jiahui and Chen, Yurui and Xu, Yueming and Huang, Ze and Zhou, Yanpeng and Yuan, Yujie and Cai, Xinyue and Huang, Guowei and Quan, Xingyue and Xu, Hang and Zhang, Li}, year={2025}, journal={arXiv preprint arXiv:2506.22242}, }

- 14D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration复旦大学数据科学学院, 华为诺亚方舟实验室 · 2025年



