EgoPro
收藏资源简介:
EgoPro is a large-scale multi-view egocentric human demonstration dataset developed by Lightwheel as part of the EgoSuite-Open100K collection, designed for Physical AI, embodied intelligence, robotics, imitation learning, and multimodal foundation model research. The dataset contains approximately 10,000 planned hours of synchronized head-view and wrist-view first-person video paired with 3D hand pose annotations, providing richer observation of human manipulation and interaction than head-view-only egocentric datasets. It consists of two major subsets: EgoProStandard, containing approximately 8,000 hours of synchronized head and wrist video with 3D hand pose annotations, and EgoProStandard-body, containing approximately 2,000 hours with additional full-body pose annotations. Temporal event-level semantic annotations are also provided as an additional annotation layer. Alternative LeRobot and MCAP representations correspond to the same recorded episodes and are not counted as additional hours. The synchronized multi-camera perspective makes EgoPro particularly suitable for studying fine-grained human-object interaction, manipulation trajectories, action understanding, imitation learning, Vision-Language-Action models, world models, robotic policy learning, and embodied agents that learn complex real-world behaviors from human demonstrations.
EgoPro is a large-scale multi-view egocentric human demonstration dataset developed by Lightwheel as part of the EgoSuite-Open100K collection, designed for Physical AI, embodied intelligence, robotics, imitation learning, and multimodal foundation model research. The dataset contains approximately 10,000 planned hours of synchronized head-view and wrist-view first-person video paired with 3D hand pose annotations, providing richer observation of human manipulation and interaction than head-view-only egocentric datasets. It consists of two major subsets: EgoProStandard, containing approximately 8,000 hours of synchronized head and wrist video with 3D hand pose annotations, and EgoProStandard-body, containing approximately 2,000 hours with additional full-body pose annotations. Temporal event-level semantic annotations are also provided as an additional annotation layer. Alternative LeRobot and MCAP representations correspond to the same recorded episodes and are not counted as additional hours. The synchronized multi-camera perspective makes EgoPro particularly suitable for studying fine-grained human-object interaction, manipulation trajectories, action understanding, imitation learning, Vision-Language-Action models, world models, robotic policy learning, and embodied agents that learn complex real-world behaviors from human demonstrations.




