xperience-10m
收藏资源简介:
Xperience-10M 是一个面向具身人工智能、机器人学、世界模型和空间智能的大规模第一人称多模态数据集,包含 1000 万次交互体验和 1 万小时的同步第一视角录制数据,涵盖六路视频流(四路鱼眼 + 两路立体)、音频、立体深度、相机位姿、双手动捕、全身动捕、IMU 以及分层语言标注(任务、子任务、动作、交互、物体),总计 28.8 亿 RGB 帧、7.2 亿深度帧、5.76 亿位姿与动捕帧、约 1 PB 数据量,是目前规模最大、结构化 3D/4D 标注最丰富的第一视角数据集之一,适用于多模态预训练、动作理解、人-物交互、机器人模仿学习、世界模型训练等研究方向,但不得用于身份识别、生物特征画像、监控或安全关键部署等超出范围的用途。
Xperience-10M is a large-scale first-person multimodal dataset tailored for embodied artificial intelligence, robotics, world modeling, and spatial intelligence. It contains 10 million interactive experiences and 10,000 hours of synchronized first-person recorded data, covering six-channel video streams (four fisheye + two stereo), audio, stereo depth, camera poses, dual-hand motion capture, full-body motion capture, IMU, and hierarchical language annotations (tasks, subtasks, actions, interactions, objects). In total, it encompasses 2.88 billion RGB frames, 720 million depth frames, 576 million pose and motion capture frames, with an overall data volume of approximately 1 PB. It stands as one of the largest first-person datasets with the most comprehensive structured 3D/4D annotations to date, supporting research areas including multimodal pre-training, action understanding, human-object interaction, robotic imitation learning, and world model training. However, its use is prohibited for out-of-scope applications such as identity recognition, biometric profiling, surveillance, or safety-critical deployments.




