混合现实空间3D人体姿态虚拟视频数据
收藏资源简介:
混合现实空间3D人体姿态虚拟视频数据聚焦复杂人机交互场景,基于 UnrealEngine 构建虚拟仿真环境,生成含 RGB 视频、深度图、3D 骨架标注的多模态数据,覆盖多角色、多场景,具备多模态融合、环境挑战覆盖全面、精度可控等优势 。在应用场景上,可深度赋能多领域:广泛服务教育培训、应急管理、安全生产与地产建筑等领域,支撑职业技能实训、安全行为监管、虚拟应急演练及沉浸式展示,为智慧教育优化教学动作示范、工业安全强化风险行为识别、数字地产实现场景沉浸式呈现 。目前已在智慧教育、工业安全、数字地产等方向落地,提供数据支撑与技术方案,从科研突破到产业实践,驱动各领域创新升级,助力教学提效、风险预警、决策优化,拓展应用边界,释放数据价值 。1.虚拟动作生成:在 UnrealEngine (UE) 环境中搭建多类虚拟场景,覆盖简化实验空间与复杂交互空间。利用骨骼驱动的人体角色模型,结合外部动作库(如 Reallusion 等),生成包含日常行为、交互动作及高难度姿态的动画序列。 2.多模态渲染与标注:在场景中布置多视角虚拟摄像机,实时渲染输出 RGB 图像与深度图 (RGB-D)。同步生成 3D 骨架关键点、相机参数及可见性标签,保证每帧均具备完整、无缺失的标注。 3. 2D 姿态推理与噪声模拟:为模拟真实应用场景中的推理偏差,对渲染生成的视频逐帧输入自研的多模态 2D 姿态估计算法。输出带噪声的 2D 关键点预测结果,用于模拟现实场景下因遮挡、光照或动作突变引入的预测抖动与误差。 4.伪深度值构建:基于深度图信息,将预测的 2D 关键点映射到三维空间,生成伪 3D 关键点坐标。该结果既保留了真实推理中的不确定性,又能作为新的信息融合使用,满足多层次算法研究需求。 5.数据存储与利用:最终数据集包含三类信息:RGB-D 视频(多视角渲染结果);高精度 3D GroundTruth(虚拟骨骼驱动生成);带噪声的 2D / 伪 3D 推理结果(模拟真实应用场景和辅助信息融合)。
This Mixed Reality Spatial 3D Human Pose Virtual Video Dataset focuses on complex human-computer interaction scenarios. It builds a virtual simulation environment based on UnrealEngine, and generates multimodal data including RGB videos, depth maps and 3D skeleton annotations. The dataset covers multiple characters and scenarios, and has advantages such as multimodal fusion, comprehensive coverage of environmental challenges, and controllable accuracy. In terms of application scenarios, it deeply empowers multiple fields: it widely serves education and training, emergency management, work safety, real estate and construction and other fields, supporting vocational skill training, safety behavior supervision, virtual emergency drills and immersive displays. It optimizes teaching action demonstrations for smart education, strengthens risk behavior recognition for industrial safety, and realizes immersive scene presentation for digital real estate. Currently, it has been implemented in fields such as smart education, industrial safety and digital real estate, providing data support and technical solutions. It drives innovative upgrades in various fields from scientific research breakthroughs to industrial practice, helps improve teaching efficiency, risk early warning and decision optimization, expands application boundaries, and releases data value. 1. Virtual Action Generation: Build various virtual scenes in the UnrealEngine (UE) environment, covering simplified experimental spaces and complex interactive spaces. Use skeleton-driven human character models combined with external motion libraries (such as Reallusion) to generate animation sequences including daily behaviors, interactive actions and high-difficulty postures. 2. Multimodal Rendering and Annotation: Arrange multi-view virtual cameras in the scene, and render and output RGB images and depth maps (RGB-D) in real time. Synchronously generate 3D skeleton key points, camera parameters and visibility tags to ensure that each frame has complete and intact annotations. 3. 2D Pose Inference and Noise Simulation: To simulate the inference deviation in real application scenarios, input the rendered video frame by frame into the independently developed multimodal 2D pose estimation algorithm. Output noisy 2D key point prediction results to simulate prediction jitter and errors caused by occlusion, lighting or abrupt motion changes in real-world scenarios. 4. Pseudo-depth Value Construction: Based on depth map information, map the predicted 2D key points to 3D space to generate pseudo-3D key point coordinates. This result not only retains the uncertainty in real inference, but can also be used as new information for fusion, meeting the needs of multi-level algorithm research. 5. Data Storage and Utilization: The final dataset includes three types of information: RGB-D videos (multi-view rendering results); high-precision 3D GroundTruth (generated by virtual skeleton-driven); noisy 2D / pseudo-3D inference results (simulating real application scenarios and assisting information fusion).




