Cubify Anything VQA (CA-VQA)
收藏资源简介:
CA-VQA数据集是基于高质量3D场景数据构建的,涵盖了各种输入信号(单一图像、度量深度图、多帧/多视图)和空间理解任务(如空间关系预测、度量大小和距离估计、3D定位)。数据集包含了丰富的输入信号,并提供了多视图图像和不同类型的度量深度图。该数据集是首个基于高质量3D地面真实数据的数据集,也是首个包含深度图(包括来自传感器的和最先进的单目估计深度)和多视图图像的数据集,覆盖了各种任务,并具有监督微调数据集和基准测试。
The CA-VQA dataset is constructed based on high-quality 3D scene data, covering diverse input signals (single images, metric depth maps, multi-frame/multi-view images) and spatial understanding tasks such as spatial relation prediction, metric size and distance estimation, and 3D localization. The dataset features rich input signals, and provides multi-view images and various types of metric depth maps. This dataset is not only the first one based on high-quality 3D ground-truth data, but also the first dataset that includes depth maps (including both sensor-derived ones and state-of-the-art monocular estimated depth) and multi-view images, covering a wide range of tasks and offering supervised fine-tuning datasets and benchmarks.




