Vision-Language-Action Instruction Tuning (VLA-IT)
收藏资源简介:
VLA-IT数据集由上海人工智能实验室创建,包含650,000个人工标注的人机交互数据,这些数据被标注了多样的指令、场景描述和问答对,并基于高质量的操控任务。数据集的创建过程采用了两阶段范式:首先进行动作预训练,然后进行视觉-语言-动作指令调整,以实现文本推理和动作生成的联合优化。VLA-IT数据集旨在解决现有视觉-语言-动作模型在任务特定操控数据上的局限性,并缓解预训练视觉-语言能力的灾难性遗忘问题。数据集的应用领域主要在机器人操控任务中,通过利用视觉-语言理解来提升操控性能,实现直观和可控的人机交互。
The VLA-IT dataset was developed by Shanghai AI Laboratory, consisting of 650,000 manually annotated human-robot interaction data samples. Each sample is annotated with diverse instructions, scene descriptions, and question-answer pairs, and is grounded in high-quality manipulation tasks. The dataset construction adopts a two-stage paradigm: first, action pre-training, followed by vision-language-action instruction tuning, to achieve joint optimization of text reasoning and action generation. The VLA-IT dataset aims to address the limitations of existing vision-language-action models on task-specific manipulation data, and mitigate catastrophic forgetting of pre-trained vision-language capabilities. Its primary application scenarios are robotic manipulation tasks, where visual-language understanding is leveraged to enhance manipulation performance and enable intuitive and controllable human-robot interaction.




