OmniAction-LIBERO
收藏资源简介:
OmniAction是一个大规模的多模态数据集,用于上下文指令跟随。它包括141,162个场景,涵盖112种技能和748种物体,包含5,096种不同的发声者音色、2,482种非语言声音事件和640种环境背景。数据集覆盖了六种上下文指令类别,包括情感线索、重叠声音、非语言线索、身份线索、二人对话和三人对话,捕捉了日常环境中的细微情感信号和复杂的多人互动。
OmniAction is a large-scale multimodal dataset designed for context-aware instruction following. It comprises 141,162 scenarios, covering 112 skills and 748 object categories, as well as 5,096 distinct speaker timbres, 2,482 non-verbal sound events, and 640 environmental backgrounds. The dataset encompasses six categories of contextual instructions, including emotional cues, overlapping sounds, non-verbal cues, identity cues, dyadic conversations, and triadic conversations, capturing subtle emotional signals and complex multi-person interactions in daily environments.
OmniAction-LIBERO 数据集概述
数据集基本信息
- 许可证: CC-BY-NC-4.0
- 任务类别: 机器人技术
- 语言: 英语
数据集状态
- 上传进度: 由于数据集规模较大,目前正在分批上传中,完整数据集将在上传完成后开放访问
数据集简介
OmniAction是一个大规模多模态上下文指令跟随数据集,专为机器人主动操作研究设计。该数据集支持从语音对话、环境声音和视觉线索中推断用户意图的新设置。
数据集规模与构成
- 总样本量: 141,162个片段
- 技能覆盖: 112种技能
- 物体数量: 748个物体
- 语音特征: 5,096种不同的说话者音色
- 声音事件: 2,482种非语言声音事件
- 环境背景: 640种环境背景
上下文指令类型
包含六类上下文指令:
- 情感线索
- 重叠语音
- 非语言线索
- 身份线索
- 二元对话
- 三元对话
数据格式
- 格式标准: RLDS(强化学习数据集标准)
- 音频处理: 按文件名排序
相关资源
- 论文: https://arxiv.org/pdf/2510.23763
- 网站: https://OpenMOSS.github.io/RoboOmni
- 模型: https://huggingface.co/fnlp/RoboOmni
- 数据集: https://huggingface.co/datasets/fnlp/OmniAction
- 代码: https://github.com/OpenMOSS/RoboOmni




