GeoGen
收藏资源简介:
GeoGen是由穆罕默德·本·扎耶德人工智能大学等机构构建的大规模几何感知生成数据集,旨在通过隐式空间世界建模增强多模态大语言模型的三维空间推理能力。该数据集包含2,241个视频和267,827个标注的三元组,数据来源于扫描的三维场景资产和互联网视频,确保了内容的广泛覆盖性与多样性。其构建过程涉及采集带有几何标注的视频数据,并生成在几何变换(如新视角合成和轨迹生成)下的交互结果。该数据集主要应用于三维视觉与机器人领域,旨在解决当前模型在空间语义与视觉感知间缺乏跨模态 grounding 的问题,通过提供几何感知的视觉反馈监督,促进模型对三维结构的直观理解与推理。
GeoGen is a large-scale geometric-aware generative dataset constructed by institutions including Mohamed bin Zayed University of Artificial Intelligence and other relevant organizations, aiming to enhance the 3D spatial reasoning capabilities of multimodal large language models through implicit spatial world modeling. This dataset contains 2,241 videos and 267,827 annotated triplets, with data sourced from scanned 3D scene assets and internet videos, ensuring broad coverage and diversity of the content. Its construction workflow includes collecting video data paired with geometric annotations, and generating interactive results under geometric transformations such as novel view synthesis and trajectory generation. This dataset is primarily utilized in the fields of 3D vision and robotics, targeting the issue that current models lack cross-modal grounding between spatial semantics and visual perception. By providing geometric-aware visual feedback supervision, it promotes the model's intuitive understanding and reasoning of 3D structures.




