MMS-VPR
收藏资源简介:
MMS-VPR是一个大型多模态数据集,用于复杂、行人专用环境中的街道级位置识别。该数据集包含78,575张注释图像和2,512个视频剪辑,跨越中国成都约70,800平方米的开放式商业区中的207个地点。每张图像都标注有精确的GPS坐标、时间戳和文本元数据,涵盖了不同的光照条件、视角和时间框架。数据集遵循系统和可复制的数据收集协议,降低了可扩展数据集创建的门槛。重要的是,该数据集形成一个固有的空间图,具有125个边缘、81个节点和1个子图,使结构感知位置识别成为可能。我们进一步定义了两个特定于应用程序的子集——Dataset_Edges和Dataset_Points,以支持细粒度和基于图的评估任务。使用传统VPR模型、图神经网络和多模态基线的广泛基准测试表明,利用多模态和结构线索可以获得显着改进。MMS-VPR促进了计算机视觉、地理空间理解和多模态推理交叉领域的未来研究。该数据集可在https://huggingface.co/datasets/Yiwei-Ou/MMS-VPR公开获取。
MMS-VPR is a large-scale multimodal dataset designed for street-level place recognition in complex, pedestrian-only environments. This dataset contains 78,575 annotated images and 2,512 video clips, spanning 207 locations across an open commercial area of approximately 70,800 square meters in Chengdu, China. Each image is annotated with precise GPS coordinates, timestamps, and textual metadata, covering diverse lighting conditions, viewpoints, and temporal frames. The dataset follows a systematic and reproducible data collection protocol, lowering the barrier to scalable dataset creation. Importantly, the dataset forms an inherent spatial graph with 125 edges, 81 nodes, and 1 subgraph, enabling structure-aware place recognition. We further define two application-specific subsets, Dataset_Edges and Dataset_Points, to support fine-grained and graph-based evaluation tasks. Extensive benchmarking using traditional VPR models, graph neural networks, and multimodal baselines demonstrates that leveraging multimodal and structural cues yields significant performance improvements. MMS-VPR facilitates future research at the intersection of computer vision, geospatial understanding, and multimodal reasoning. This dataset is publicly available at https://huggingface.co/datasets/Yiwei-Ou/MMS-VPR.

- 1MMS-VPR: Multimodal Street-Level Visual Place Recognition Dataset and Benchmark奥克兰大学 · 2025年



