遇见数据集

ChikaKomari/26summe-camp-10-data

收藏
Hugging Face2026-05-21 更新2026-05-31 收录
官方服务:

资源简介:

StarVLA数据集是一个用于视觉语言动作(VLA)训练的多模态机器人数据集集合,包含多个子数据集:CALVIN(基于Panda机器人的任务数据集,包含17870个episode和1071743帧,用于动作预测)、OXE(包括Bridge/WidowX和Fractal/RT-1数据集,分别涉及widowx和google_robot机器人,用于跨实施例训练)、RoboCasa(基于GR1/Fourier Hands的24个任务数据集,每个任务1000个episode,用于桌面操作任务)和RoboTwin 2.0(包含Clean和Randomized任务,涉及aloha机器人,用于双臂操作任务)。这些数据集覆盖了不同机器人平台、任务类型和动作模态,支持多框架训练(如QwenPI、QwenOFT等),用于训练视觉语言动作模型以实现机器人控制。

The StarVLA dataset is a collection of multimodal robotic datasets for Vision-Language-Action (VLA) training, comprising several sub-datasets: CALVIN, a task dataset based on the Panda robot containing 17,870 episodes and 1,071,743 frames for action prediction; OXE, which includes the Bridge/WidowX and Fractal/RT-1 datasets involving the widowx and google_robot robots respectively and used for cross-embodiment training; RoboCasa, a dataset of 24 tasks based on GR1/Fourier Hands with 1,000 episodes per task targeting tabletop manipulation tasks; and RoboTwin 2.0, which includes Clean and Randomized tasks involving the Aloha robot and designed for dual-arm manipulation tasks. These datasets cover diverse robotic platforms, task types and action modalities, support training with multiple frameworks (e.g., QwenPI, QwenOFT, etc.), and are used to train Vision-Language-Action models for robotic control.

提供机构:
ChikaKomari
二维码
社区交流群
二维码
科研交流群
商业服务