GIQ
收藏资源简介:
GIQ数据集是一套旨在评估视觉和视觉-语言基础模型几何推理能力的综合基准。该数据集包含了224个多样化的多面体图像,包括柏拉图、阿基米德、约翰逊和卡塔兰固体,以及星形和复合形状,涵盖了不同的复杂性和对称性。数据集包括模拟和真实世界图像,从多个视角捕获,以评估模型对对称性的识别、从单张图像中重建复杂几何形状的能力,以及在不同视角和真实世界条件下准确推理形状等价性的能力。GIQ数据集的创建过程涉及使用Mitsuba物理渲染器生成模拟多面体,以及从纸张构建物理模型并在各种室内外环境中拍摄。该数据集为诊断和提升视觉系统中的几何智能提供了一个有针对性的基准,为未来改进空间感知和3D感知视觉推理奠定了基础。
The GIQ dataset is a comprehensive benchmark designed to evaluate the geometric reasoning capabilities of visual and vision-language foundation models. This dataset includes 224 diverse polyhedral images, covering Platonic, Archimedean, Johnson, and Catalan solids, as well as star-shaped and compound shapes, which span varying levels of complexity and symmetry. The dataset comprises both simulated and real-world images captured from multiple viewpoints, to assess models' abilities to recognize symmetry, reconstruct complex geometric shapes from a single image, and accurately reason about shape equivalence across different viewpoints and real-world conditions. The creation of the GIQ dataset involved using the Mitsuba physically based renderer to generate simulated polyhedra, as well as constructing physical models from paper and photographing them across various indoor and outdoor environments. This dataset provides a targeted benchmark for diagnosing and enhancing geometric intelligence in visual systems, laying a foundation for future improvements to spatial awareness and 3D-aware visual reasoning.




