DepthVLM-Bench
收藏资源简介:
DepthVLM-Bench是一个专为视觉语言模型设计的统一室内外度量深度估计基准。它提供了多样化的室内和室外场景,并以统一的、与视觉语言模型兼容的格式提供了度量深度标注。其核心目标是使大规模多模态模型能够联合学习密集的几何(深度)预测与多模态理解能力。数据集的主要特点包括:支持统一的室内和室外环境度量深度估计;数据格式专门适配视觉语言模型;为多模态基础模型提供密集的深度监督信号;整体设计旨在支持可扩展的多模态训练。该数据集与题为《Unlocking Dense Metric Depth Estimation in VLMs》的研究论文相关联。
DepthVLM-Bench is a unified indoor-outdoor metric depth estimation benchmark designed for vision-language models. It provides diverse indoor and outdoor scenes with metric depth annotations in a unified format compatible with vision-language models. The core goal is to enable large-scale multimodal models to jointly learn dense geometric (depth) prediction and multimodal understanding capabilities. Key features of the dataset include: support for unified indoor and outdoor metric depth estimation; data format specifically adapted for vision-language models; provision of dense depth supervision signals for multimodal foundation models; and overall design aimed at supporting scalable multimodal training. This dataset is associated with the research paper titled Unlocking Dense Metric Depth Estimation in VLMs.
数据集概述
DepthVLM-Bench 是一个为视觉语言模型(VLM)设计的统一室内外度量深度估计基准数据集。
核心特征
- 统一室内外场景:包含多样化的室内和室外场景,并提供度量深度标注。
- VLM兼容格式:数据采用统一的、与视觉语言模型兼容的格式。
- 密集深度监督:提供密集的深度信息,用于多模态基础模型的训练。
- 可扩展的多模态训练:专为大规模多模态联合学习而设计。
相关论文
- 论文标题:Unlocking Dense Metric Depth Estimation in VLMs
- 论文地址:https://arxiv.org/abs/2605.15876
使用方式
官方代码仓库(https://github.com/hanxunyu/DepthVLM)提供了数据处理、评估脚本和可视化示例等资源。
引用信息
bibtex @article{yu2026unlocking, title={Unlocking Dense Metric Depth Estimation in VLMs}, author={Hanxun Yu and Xuan Qu and Yuxin Wang and Jianke Zhu and Lei Ke}, journal={arXiv preprint arXiv:2605.15876}, year={2026} }
许可信息
- 开源许可:Apache-2.0





