OmniAction
收藏资源简介:
OmniAction是一个大规模的多模态数据集,用于上下文指令的跟随。它包含了141,162个场景,涵盖了112种技能和748个对象,并丰富了5,096种不同的说话人声音、2,482种非语言声音事件和640种环境背景。数据集覆盖了六种类型的上下文指令,包括情感线索、重叠声音、非语言线索、身份线索、二元对话和三元对话,捕捉了日常环境中的微妙情感信号和复杂的多人互动。
OmniAction is a large-scale multimodal dataset developed for following contextual instructions. It includes 141,162 scenarios, covering 112 skills and 748 objects, as well as 5,096 distinct speaker voices, 2,482 non-verbal sound events, and 640 environmental backgrounds. The dataset covers six types of contextual instructions, namely emotional cues, overlapping speech, non-verbal cues, identity cues, dyadic conversations, and triadic conversations, capturing subtle emotional signals and complex multi-person interactions in daily environments.
OmniAction 数据集概述
数据集基本信息
- 许可证: CC-BY-NC-4.0
- 任务类别: 机器人技术
- 语言: 英语
数据集状态
- 由于数据集规模较大,目前正在分批上传
- 完整数据集将在上传完成后开放访问
数据集规模与内容
- 总样本量: 141,162个片段
- 技能覆盖: 112种技能
- 物体覆盖: 748个物体
- 语音多样性: 5,096种不同的说话者音色
- 非语言声音事件: 2,482种
- 环境背景: 640种
上下文指令类型
数据集涵盖六类上下文指令:
- 情感线索
- 重叠语音
- 非语言线索
- 身份线索
- 二元对话
- 三元对话
技术规格
- 数据格式: RLDS(强化学习数据集标准)
- 音频处理: 按文件名排序
相关资源
- 论文: https://arxiv.org/pdf/2510.23763
- 项目网站: https://OpenMOSS.github.io/RoboOmni
- 模型: https://huggingface.co/fnlp/RoboOmni
- 代码库: https://github.com/OpenMOSS/RoboOmni
引用信息
bibtex @article{wang2025roboomni, title={RoboOmni: Proactive Robot Manipulation in Omni-modal Context}, author={Siyin Wang and Jinlan Fu and Feihong Liu and Xinzhe He and Huangxuan Wu and Junhao Shi and Kexin Huang and Zhaoye Fei and Jingjing Gong and Zuxuan Wu and Yugang Jiang and See-Kiong Ng and Tat-Seng Chua and Xipeng Qiu}, journal={arXiv preprint arXiv:2510.23763}, year={2025}, url={https://arxiv.org/abs/2510.23763}, archivePrefix={arXiv}, primaryClass={cs.RO}, }




