EarthWhere
收藏资源简介:
EarthWhere是一个全面的地标定位基准,用于评估视觉语言模型(VLMs)的地标定位能力。该数据集包含了810张全球分布的图像,跨越两个互补的地标定位规模:WhereCountry和WhereStreet。WhereCountry任务包含500个多选题,要求模型识别国家级别的地标;WhereStreet任务包含310个需要多步推理的精细街道级别识别任务。数据集的创建过程涉及了对图像进行人工验证的关键视觉线索的收集和推理过程的评估。EarthWhere旨在解决视觉语言模型在开放世界条件下地标定位的问题,这对于现实生活中的搜索和救援、城市规划或环境监测等领域至关重要。
EarthWhere is a comprehensive landmark localization benchmark for evaluating the landmark localization capabilities of Vision-Language Models (VLMs). This dataset contains 810 globally distributed images, spanning two complementary landmark localization scales: WhereCountry and WhereStreet. The WhereCountry task consists of 500 multiple-choice questions that require models to recognize national-level landmarks; the WhereStreet task includes 310 fine-grained street-level recognition tasks that demand multi-step reasoning. The dataset creation process entails collecting key visual cues from images with manual verification, as well as evaluating the associated reasoning procedures. EarthWhere aims to address the challenge of landmark localization for vision-language models under open-world conditions, which is critical for real-world applications such as search and rescue, urban planning, and environmental monitoring.

- 1通过美国加州大学圣克鲁斯分校, 美国佛罗里达中央大学, 美国哥伦比亚大学, 亚马逊研究院, 美国加州大学伯克利分校 · 2025年



