GeoMapSet: A POI-driven multimodal dataset for map image understanding
收藏资源简介:
Maps are essential yet complex media for encoding geospatial information, but the scarcity of large-scale, semantically rich map–text paired data limits the training and evaluation of multimodal models for map understanding. We introduce GeoMapSet, a large-scale multimodal benchmark designed to systematically evaluate VLM performance on map understanding. Constructed via an automated pipeline using OpenStreetMap (OSM) as a unified source, GeoMapSet comprises 21, 870 high-resolution map images from 7 cities across 4 continents with 5 distinct map annotation languages (Chinese, English, French, Japanese, Korean), paired with structured textual descriptions in English, covering POI quantity, functional density/diversity, and spatial distribution patterns. A two-level category hierarchy (22 fine-grained icon types→6 broad functional classes) bridges visual icon diversity with linguistic semantic compactness.



