RoVI Book
收藏资源简介:
RoVI Book数据集由上海人工智能实验室的研究人员创建,包含15,000个图像-文本问答对,旨在帮助视觉语言模型学习理解RoVI(机器人视觉指令)的能力。数据集覆盖了64%的单步任务和36%的多步任务,涉及移动物体、旋转物体、拾取、打开/关闭抽屉/橱柜等五种基本的操作技能。数据集提供了RoVI分析、任务名称、细粒度规划步骤和Python函数的答案,通过GPT-4o生成并经过语义过滤。数据集的创建是为了解决自然语言在机器人任务定义中的空间精度不足问题,通过手绘的符号表示来传达更精确的空间时间信息,使机器人能够更好地理解RoVI并执行精确的动作。
The RoVI Book dataset was developed by researchers from the Shanghai AI Laboratory, comprising 15,000 image-text question-answer pairs. It is designed to assist vision-language models in acquiring the ability to understand RoVI (Robot Visual Instructions). The dataset covers 64% of single-step tasks and 36% of multi-step tasks, involving five core manipulation skills: moving objects, rotating objects, picking, opening and closing drawers and cabinets. The dataset provides answers including RoVI analysis, task names, fine-grained planning steps, and Python functions, which were generated using GPT-4o and filtered via semantic checks. This dataset was created to address the issue of insufficient spatial precision of natural language in robot task definition. By adopting hand-drawn symbolic representations to convey more precise spatiotemporal information, it enables robots to better comprehend RoVI and execute precise actions.

- 1Robotic Visual Instruction上海人工智能实验室 · 2025年



