步入式混合现实空间2D人体姿态深度图数据
收藏资源简介:
该数据面向复杂人机交互与高精度动作识别,广泛应用于虚拟现实、增强现实、智能制造、教育培训、应急管理、安全生产及地产建筑等领域。通过高质量深度与RGB数据融合,可在复杂光照、遮挡及多人交互环境中实现精准人体姿态捕捉,支持裸眼3D沉浸式教学、虚拟实验、应急演练、实时监控与建筑展示,有效提升交互体验与应用效率。1.数据采集:通过自主研发的深度成像设备,在混合现实场景中持续采集多帧深度图,重点覆盖光照复杂、阴影频繁和低光等环境。采集过程严格遵守法规与安全规范,确保数据 合规可靠,同时避免泄露个人隐私,为后续姿态识别算法提供稳定输入。 2. 人体检测:采用针对深度图优化的自训练目标检测算法,对每一帧深度数据进行人体区 域定位。该算法结合了大规模真实场景深度样本进行训练,能够在复杂光影与背景变化条 件下保持高鲁棒性。算法输出人体外接矩形框坐标及其置信度,确保后续姿态估计步骤的 数据基础精准可靠。 3. 关键点提取:在检测到的人体区域基础上,引入多模型集成的人体姿态估计算法。输入为人体 ROI 的深度图像,输出为人体 17 个关键点(如肩、肘、腕、髋等)的空间像素坐标及置信度。由于深度图具有更强的几何约束,模型能够在复杂环境中实现稳定、精确的骨骼关键点定位。 4. 动态指标计算与误差识别:基于连续帧的推理结果,对各个关节在时间维度上的运动变化(速度和加速度)进行计算。同时结合关键点的置信度分布,对可能存在预测偏差的帧进行标记,从而为后续人工审核提供依据。 5. 人工修订:对于运动加速度超过设定阈值的异常帧,或关键点置信度不足的低质量预测,由专业人员进行人工校正。在人工修订过程中,常见的关键点预测误差类型主要包括: Jitter:预测点在真实位置附近存在微小抖动;Miss:预测点与真实位置相距较远,未能落在相应部位邻近区域; Inversion:同一人体实例内部出现混淆,例如将左肘识别为右肘;Swap:不同人体实例之间的混淆,预测点被错误地落在另一人的对应部位附近。 6. 时空平滑优化:利用自研的时空神经网络,对经过修正的关键点序列开展平滑处理。该模型综合考虑时间连续性与空间一致性,有效减少因采集噪声或检测误差带来的抖动,使最终输出的深度姿态序列更加自然、流畅,显著提升数据的整体质量。
This dataset is targeted at complex human-computer interaction and high-precision action recognition, and is widely applied in fields such as virtual reality (VR), augmented reality (AR), intelligent manufacturing, education and training, emergency management, work safety, and real estate and construction. Through high-quality fusion of depth and RGB data, it can achieve accurate human pose capture in complex lighting, occlusion and multi-person interaction environments, supporting autostereoscopic 3D immersive teaching, virtual experiments, emergency drills, real-time monitoring and architectural display, effectively improving interaction experience and application efficiency. 1. Data Collection: Using independently developed depth imaging equipment, multiple frames of depth maps are continuously collected in mixed reality scenarios, with key coverage of environments with complex lighting, frequent shadows and low light. The collection process strictly complies with laws, regulations and safety norms to ensure data compliance and reliability, while avoiding leakage of personal privacy, providing stable input for subsequent pose recognition algorithms. 2. Human Detection: A self-trained object detection algorithm optimized for depth maps is adopted to locate human regions in each frame of depth data. Trained with large-scale real-world depth samples, this algorithm maintains high robustness under complex lighting and background changes. The algorithm outputs the coordinates of human bounding boxes and their confidence levels, ensuring the accurate and reliable data basis for the subsequent pose estimation step. 3. Key Point Extraction: Based on the detected human regions, a multi-model ensemble human pose estimation algorithm is introduced. The input is the depth image of the human ROI (Region of Interest), and the output is the spatial pixel coordinates and confidence levels of 17 human key points (such as shoulders, elbows, wrists, hips, etc.). Since depth maps have stronger geometric constraints, the model can achieve stable and accurate skeletal key point positioning in complex environments. 4. Dynamic Index Calculation and Error Identification: Based on the inference results of consecutive frames, the temporal motion changes (velocity and acceleration) of each joint are calculated. Meanwhile, combined with the confidence distribution of key points, frames that may have prediction deviations are marked, providing a basis for subsequent manual review. 5. Manual Revision: For abnormal frames where the motion acceleration exceeds the set threshold, or low-quality predictions with insufficient key point confidence, professional personnel perform manual correction. During the manual revision process, the common types of key point prediction errors mainly include: - Jitter: The predicted point has slight jitters near the real position; - Miss: The predicted point is far from the real position and fails to fall in the adjacent area of the corresponding body part; - Inversion: Confusion occurs within the same human instance, for example, identifying the left elbow as the right elbow; - Swap: Confusion between different human instances, where the predicted point is incorrectly located near the corresponding part of another person. 6. Spatiotemporal Smoothing Optimization: Using independently developed spatiotemporal neural networks, smoothing processing is carried out on the corrected key point sequence. This model comprehensively considers temporal continuity and spatial consistency, effectively reducing jitters caused by collection noise or detection errors, making the final output depth pose sequence more natural and smooth, and significantly improving the overall quality of the dataset.




