RoVid-X
收藏资源简介:
RoVid-X是一个大规模机器人视频生成数据集,包含4M机器人视频片段(超过10,000小时),涵盖1300+细粒度机器人技能,支持多样化的动作和任务原语。数据集提供多模态物理注释,包括RGB、深度和光流信息,覆盖多种机器人类型、场景和动作技能,并包含丰富的物体交互,以实现复杂和真实的机器人行为建模。
RoVid-X is a large-scale robotic video generation dataset. It encompasses 4 million robotic video clips (totaling over 10,000 hours), covering more than 1,300 fine-grained robotic skills and supporting diverse motion and task primitives. The dataset provides multimodal physical annotations including RGB, depth, and optical flow information. It covers a broad spectrum of robot types, scenarios, motion skills, and incorporates rich object interactions, enabling the modeling of complex and realistic robotic behaviors.
RoVid-X 数据集概述
基本信息
- 数据集名称:RoVid-X
- 发布机构:DAGroup-PKU
- 语言:英语
- 许可协议:CC-BY-4.0
- 规模分类:大于1TB
- 任务类别:图像到视频生成
- 标签:机器人视频生成、文本到视频、图像到视频、视频生成、大规模、基准测试、评估
核心特性
- 规模:包含400万个机器人视频片段,总计超过1万小时,适用于大规模视频生成训练。
- 技能覆盖:涵盖1300多种细粒度机器人技能,涉及多样化的动作和任务原语。
- 多模态物理标注:提供RGB、深度和光流等多模态物理标注信息。
- 多样性:涵盖多种机器人类型、场景和动作技能,具有多机器人和多任务多样性。
- 对象交互:包含丰富的对象交互内容,支持复杂且真实的机器人行为建模。
数据结构
数据集以JSON格式为每个视频片段提供结构化标注,每个条目通过视频文件名进行索引。标注内容包括动词、任务描述、简短描述和详细描述。
下载方式
可通过Hugging Face官方CLI工具直接下载数据集。中国大陆用户可使用镜像地址加速下载。
引用信息
如果使用本数据集,请引用相关论文:
@misc{deng2026rethinkingvideogenerationmodel, title={Rethinking Video Generation Model for the Embodied World}, author={Yufan Deng and Zilin Pan and Hongyu Zhang and Xiaojie Li and Ruoqing Hu and Yufei Ding and Yiming Zou and Yan Zeng and Daquan Zhou}, year={2026}, eprint={2601.15282}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2601.15282}, }




