TUM/EgoExOR
收藏资源简介:
EgoExOR-HQ是一个用于手术活动理解的自中心-外中心手术室数据集。该数据集旨在解决手术室中快速、遮挡严重环境下外科医生、护士和设备之间精确协调的需求,为提升安全性和效率提供先进的感知模型基础。现有数据集通常只提供部分自中心视角或稀疏的外中心多视角上下文,而EgoExOR首次融合了第一人称和第三人称视角。数据集覆盖了94分钟(以15 FPS采样,共84,553帧)的两种模拟脊柱手术过程:超声引导针插入和微创脊柱手术。它集成了多种模态数据:自中心视角包括来自可穿戴眼镜的RGB视频、注视跟踪、手部跟踪和音频;外中心视角包括来自RGB-D相机的RGB和深度图像以及超声图像。此外,数据集提供了丰富的标注信息,包括36个实体和22种关系(共568,235个三元组),用于场景图生成任务。该数据集为手术室感知研究提供了多模态、时间同步的资源,支持细粒度视觉分析和跨模态相关性研究。
EgoExOR-HQ is an ego-exo-centric operating room dataset for surgical activity understanding. Operating rooms demand precise coordination among surgeons, nurses, and equipment in a fast-paced, occlusion-heavy environment, necessitating advanced perception models to enhance safety and efficiency. Existing datasets either provide partial egocentric views or sparse exocentric multi-view context, but do not explore the comprehensive combination of both. EgoExOR is the first OR dataset and accompanying benchmark to fuse first-person and third-person perspectives. Spanning 94 minutes (84,553 frames at 15 FPS) of two emulated spine procedures—Ultrasound-Guided Needle Insertion and Minimally Invasive Spine Surgery—EgoExOR integrates egocentric data (RGB, gaze, hand tracking, audio from wearable glasses) and exocentric data (RGB and depth from RGB-D cameras, ultrasound imagery). It also includes annotations with 36 entities and 22 relations (568,235 triplets) for scene graph generation. This dataset sets a new foundation for OR perception, offering a rich, multimodal resource for next-generation clinical perception.




