RoboSPA
收藏资源简介:
RoboSPA是一个大规模机器人操作数据集与基准,由浙江大学等机构联合构建,旨在系统性评估视觉-语言-动作(VLA)模型的具身推理能力。该数据集包含527K条轨迹和997小时视频,覆盖细粒度空间推理与长程程序规划两大核心维度,共计10个任务类别、56个基础任务,每个任务按5个难度等级扩展为280个变体,并采用5种机器人本体执行,场景涵盖干净与领域随机化环境。数据集基于SAPIEN模拟器和RoboTwin 2.0框架构建,由10名研究生设计任务并经过双重审核,确保了质量与多样性。RoboSPA专注于诊断VLA模型在复杂空间关系、记忆依赖执行及多步程序规划中的瓶颈,为开发更可靠、可泛化的具身智能体提供了具有挑战性的测试平台。
RoboSPA is a large-scale robotic manipulation dataset and benchmark jointly constructed by Zhejiang University and other institutions, aiming to systematically evaluate the embodied reasoning capabilities of vision-language-action (VLA) models. This dataset contains 527K trajectories and 997 hours of video, covering two core dimensions: fine-grained spatial reasoning and long-range procedural planning. It includes a total of 10 task categories and 56 basic tasks, with each basic task expanded into 280 variants across 5 difficulty levels, and is implemented using 5 types of robot embodiments. The scenarios cover both clean environments and domain-randomized settings. Built on the SAPIEN simulator and the RoboTwin 2.0 framework, the dataset's tasks were designed by 10 graduate students and underwent double reviews to ensure quality and diversity. RoboSPA focuses on diagnosing the bottlenecks of VLA models in complex spatial relationships, memory-dependent execution and multi-step procedural planning, providing a challenging testbed for developing more reliable and generalizable embodied AI agents.
数据集概述:RoboSPA
当前状态: 该数据集及配套代码尚未正式发布,目前处于准备阶段。项目页面仅发布了预告声明,表示相关内容“即将推出”,并邀请用户持续关注。
内容说明: 由于数据集文件和代码尚未公开,当前无法提供具体的数据构成、规模、标注格式、应用场景或使用方式等详细信息。页面内未包含任何实质性数据描述或示例。
访问提示: 项目的主页地址为 https://github.com/fanzhenxuan/RoboSPA,但该地址目前仅承载了上述预告信息。建议有意使用者关注该页面的后续更新,以获取正式发布的数据集和代码。

- 1RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?浙江大学; 电子科技大学; 华南师范大学 · 2026年



