EgoHOS
收藏资源简介:
EgoHOS数据集由宾夕法尼亚大学创建,包含11,243张第一人称视角的图像,每张图像都有精细的手部和交互对象的像素级分割标签。该数据集涵盖了多种日常活动,从近1000个视频中抽样得到,包括Ego4D、EPIC-KITCHEN、THU-READ等。数据集的创建过程中,研究团队采用了上下文感知合成数据增强技术,以适应分布外的YouTube第一人称视频。EgoHOS数据集的应用领域广泛,包括手部状态分类、视频活动识别、手部-对象交互的3D网格重建以及第一人称视频中的手部透明化等,旨在解决复杂场景下的手部和对象交互理解问题。
The EgoHOS dataset was developed by the University of Pennsylvania. It comprises 11,243 first-person images, each paired with fine-grained pixel-level segmentation labels for hands and their interacting objects. Covering diverse daily activities, the dataset is sampled from nearly 1,000 videos sourced from existing datasets including Ego4D, EPIC-KITCHEN, THU-READ, and others. During its construction, the research team adopted context-aware synthetic data augmentation techniques to adapt to out-of-distribution YouTube first-person videos. The EgoHOS dataset has broad application scenarios, such as hand state classification, video activity recognition, 3D mesh reconstruction of hand-object interactions, and hand transparency in first-person videos, among others, aiming to address the challenge of understanding hand and object interactions in complex real-world scenes.




