SynthHuman
收藏资源简介:
SynthHuman数据集是由微软剑桥团队创建的一个高保真合成数据集,旨在解决人本视觉任务中数据多样性和标注质量的问题。数据集包含30万张分辨率为384x512的图像,涵盖了人脸、上半身和全身场景,并提供了软前景掩码、表面法线和深度标注。该数据集在相对深度估计、表面法线估计和软前景分割等密集预测任务上取得了最先进的准确性,同时模型训练和推理的成本仅为类似精度的基础模型的很小一部分。
The SynthHuman dataset is a high-fidelity synthetic dataset developed by the Microsoft Cambridge Team, which aims to address the challenges of data diversity and annotation quality in human-centric visual tasks. It contains 300,000 images with a resolution of 384×512, covering facial, upper-body, and full-body scenarios, and provides soft foreground masks, surface normals, and depth annotations. This dataset achieves state-of-the-art accuracy on dense prediction tasks including relative depth estimation, surface normal estimation, and soft foreground segmentation, while the cost of model training and inference is only a small fraction of that of baseline models with comparable precision.
DAViD 数据集概述
基本信息
- 数据集名称: DAViD (Data-efficient and Accurate Vision Models from Synthetic Data)
- 发布会议: International Conference on Computer Vision 2025
- 作者: Fatemeh Saleh, Sadegh Aliakbarian, Charlie Hewitt, Lohit Petikam, Xiao-Xian, Antonio Criminisi, Thomas J. Cashman, Tadas Baltrušaitis
- 相关资源: 论文 | arXiv | 视频 | 数据集与模型
数据集描述
- 数据来源: 完全使用合成数据训练模型
- 合成数据管道: 基于Hewitt等人的数据生成流程,结合Petikam等人的更新面部模型
- 数据集名称: SynthHuman
- 数据规模: 300K张图像
- 图像分辨率: 384×512
- 内容覆盖: 面部、上半身和全身场景,比例均等
- 多样性设计: 包含多样化的姿势、环境、光照和外观
- 标注信息: 每张图像包含软前景掩码、表面法线和深度真实标注
模型架构
- 基础架构: 密集预测变换器(DPT)的变体
- 特点:
- 支持可变输入分辨率
- 单一模型架构处理三个密集预测任务
- 多任务学习能力
性能表现
- 推理速度: 低至21毫秒/帧(在NVIDIA A100上的大型多任务模型)
- 任务表现:
- 深度估计
- 表面法线估计
- 软前景分割
- 优势:
- 高精度
- 高效训练和推理
- 良好的泛化能力
引用信息
bibtex @misc{saleh2025david, title={{DAViD}: Data-efficient and Accurate Vision Models from Synthetic Data}, author={Fatemeh Saleh and Sadegh Aliakbarian and Charlie Hewitt and Lohit Petikam and Xiao-Xian and Antonio Criminisi and Thomas J. Cashman and Tadas Baltrušaitis}, year={2025}, eprint={2507.15365}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2507.15365}, }
机构信息
- 研究机构: Mixed Reality & AI Lab - Cambridge
- 版权信息: © Microsoft 2025




