ERGeoBench (Embodied Reasoning Geo-localization Benchmark)
收藏资源简介:
ERGeoBench是由北京邮电大学等机构联合构建的首个面向实体化推理与地理定位的多模态大模型诊断基准。该数据集包含全球范围内精心采集的2,207个高分辨率街景全景图,通过可控相机模型转化为以自我为中心的感知环境,支持智能体通过偏航、俯仰和缩放动作主动获取观测视角。数据构建过程融合了大规模地理数据收集与实体化视角合成技术,形成了涵盖单视图、全景视图和实体化视图的三层评估体系。该数据集旨在系统评估多模态大模型在实体化地理定位任务中的四大核心能力——基础感知、空间意识、常识推理与地理定位推理,为解决传统静态地理定位方法在动态交互与证据细化方面的局限性提供标准化测试平台。
ERGeoBench is the first multimodal large language model diagnostic benchmark for embodied reasoning and geolocation, jointly constructed by Beijing University of Posts and Telecommunications and other institutions. This dataset contains 2,207 high-resolution street view panoramas carefully collected worldwide, which are converted into an egocentric perceptual environment via a controllable camera model, enabling AI Agents to actively acquire observation perspectives through yaw, pitch, and zoom movements. The dataset construction process integrates large-scale geographic data collection and embodied view synthesis technology, forming a three-tier evaluation system covering single-view, panoramic view, and embodied view. This dataset aims to systematically evaluate four core capabilities of multimodal large language models in embodied geolocation tasks: basic perception, spatial awareness, commonsense reasoning, and geolocation reasoning, providing a standardized test platform to address the limitations of traditional static geolocation methods in dynamic interaction and evidence refinement.

- 1ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models北京邮电大学·网络与交换技术国家重点实验室; 上海交通大学·材料科学与工程学院; 中国移动研究院; 南洋理工大学·计算与数据科学学院 · 2026年



