LA4-33K
收藏资源简介:
LA4-33K 是一个由上海交通大学与阿里巴巴集团等机构联合构建的语言-动作对数据集,旨在为视觉-语言-动作模型提供去视觉化的语言-动作监督信号。该数据集包含 33,000 条语言-动作片段,其内容源于对现有专家演示轨迹的分解与重组,无需额外采集机器人数据,每个片段均将原子化的低层动作描述与对应的动作序列精确配对。数据集的创建过程是通过将完整的演示轨迹拆解为可重用的原子操作单元(如移动、抓取、旋转等),并自动生成相应的语言描述,从而显式地构建语言与动作间的映射关系。该数据集主要应用于机器人操作技能的预训练领域,其核心目标是解决传统视觉-语言-动作联合训练中语言监督信号稀疏、模型过度依赖视觉捷径的问题,通过强化语言对动作执行的条件化先验知识,提升策略的泛化能力和在视觉扰动下的鲁棒性。
LA4-33K is a language-action pair dataset jointly constructed by Shanghai Jiao Tong University, Alibaba Group and other institutions, aiming to provide vision-agnostic language-action supervision signals for vision-language-action models. This dataset contains 33,000 language-action segments, whose content is derived from the decomposition and recombination of existing expert demonstration trajectories, without requiring additional robot data collection. Each segment accurately pairs atomic low-level action descriptions with their corresponding action sequences. The dataset is created by decomposing complete demonstration trajectories into reusable atomic operation units such as moving, grasping, rotating, etc., and automatically generating corresponding language descriptions, thereby explicitly constructing the mapping relationship between language and actions. This dataset is mainly applied in the field of robotic manipulation skill pre-training. Its core goal is to address the issues of sparse language supervision signals and the model's over-reliance on visual shortcuts in traditional vision-language-action joint training. By strengthening the conditional prior knowledge of language for action execution, it enhances the generalization ability of the policy and its robustness under visual perturbations.
LA4VLA 数据集
数据集概述
- 名称:LA4VLA
- 来源:MINT-SJTU 研究团队
- 托管平台:GitHub
用途
该数据集用于视觉-语言-动作(VLA)相关研究,具体定义和用途需参考 GitHub 仓库详情。
访问地址
- GitHub 仓库:https://github.com/MINT-SJTU/LA4VLA
注:该数据集详情页README内容仅为标题,未提供更详细的描述信息。

- 1LA4VLA: Learning to Act without Seeing via Language-Action Pretraining上海交通大学·人工智能学院; 阿里巴巴集团; 南洋理工大学; 阿卜杜拉国王科技大学 · 2026年



