EgoSPT
收藏资源简介:
EgoSPT是由密歇根州立大学与英伟达研究院联合创建的具身智能空间提示操作数据集,旨在研究基于首帧空间提示的视觉轨迹预测问题。该数据集包含2841个抓放操作视频,每个视频约5秒时长并降采样至10fps,通过改进的通用操作接口采集,整合了GoPro鱼眼镜头视频与iPhone SLAM系统恢复的6自由度末端执行器轨迹数据。数据构建过程由九名专家完成,涵盖五个视觉相似叉子与九类目标容器的组合,并设计三个渐进式场景以评估模型分布内性能与跨场景泛化能力。该数据集主要应用于机器人操作策略学习领域,通过提供首帧物体与目标的空间标注、自我中心视觉观察和精确运动轨迹,解决在杂乱环境中基于稀疏空间意图生成时序运动规划的核心挑战。
EgoSPT is an embodied intelligence spatial prompting manipulation dataset jointly created by Michigan State University and NVIDIA Research, which aims to study the visual trajectory prediction problem based on first-frame spatial prompts. This dataset contains 2841 grasp-and-place operation videos, each approximately 5 seconds long and downsampled to 10fps. Collected via an improved general manipulation interface, it integrates GoPro fisheye lens footage and 6-degree-of-freedom (6DoF) end-effector trajectory data recovered by the iPhone SLAM system. The dataset construction process was completed by nine experts, covering combinations of five visually similar forks and nine types of target containers. Three progressive scenarios were designed to evaluate the model's in-distribution performance and cross-scenario generalization capability. Mainly applied in the field of robot manipulation policy learning, this dataset provides spatial annotations of first-frame objects and targets, egocentric visual observations and precise motion trajectories, to address the core challenge of generating temporal motion planning based on sparse spatial intentions in cluttered environments.
根据您提供的HTML内容,以下是该数据集详情页面的关键信息总结。
数据集概述
数据集名称:EgoSPT
所属项目:SP-VTP(Spatially Prompted Visual Trajectory Prediction)
发布时间:2026年(预印本)
发布机构:密歇根州立大学 & NVIDIA Research
核心任务
SP-VTP(空间提示的视觉轨迹预测):给定第一帧图像中的物体和目标空间提示(bounding box),模型从流式自我中心观察中预测未来的相对末端执行器轨迹。
数据集规模
- 总片段数:2,841 个自我中心操作片段
- 轨迹样本数:112,856 个
- 场景分割:3 个场景级划分(训练/验证/测试)
数据采集
- 采集设备:改进的 Universal Manipulation Interface (UMI)
- 数据内容:每个片段包含自我中心视频、初始帧中物体和目标的边界框标注、恢复的末端执行器姿态
- 验证协议:场景感知验证协议,将相关场景单元排除在训练集外,评估跨场景泛化能力
标注与处理
- 标注方式:轻量级标注流程,在第一帧中标注物体和目标边界框
- 提供工具:标注工具(记录两个边界框)和修改工具(支持人工检查和修正)
- 数据可视化:每个处理后的片段同步显示自我中心视频、恢复的姿态运动和夹爪宽度信号
使用说明
- 数据下载:页面提供数据下载链接(具体地址未在HTML中列出)
- 论文引用:提供 BibTeX 引用格式
- 代码:代码即将发布(链接标记为“coming soon”)




