Evaluating Prompting Approaches for Indoor Navigation with MLLMs
收藏资源简介:
Open data results for: The Advantages of Simplicity: Evaluating Prompting Approaches for Indoor Navigation with MLLMsThis paper investigates the capability of multimodal large language models to interpret indoor floorplans and generate human‑like navigation instructions without relying on explicit graph extraction or symbolic spatial representations. We evaluate several prompting strategies, including few‑shot prompting, self‑verification, and multi‑step reasoning, across three real-world environments. A standardized set of navigation queries is used to assess the model’s ability to infer spatial layout, identify points of interest, and produce coherent step‑by‑step directions. Results show that while MLLMs can reliably identify high‑level routes, their performance is significantly affected by map complexity, visual clutter, and orientation ambiguities. Graph‑style map simplification improves accuracy, but persistent errors in left–right reasoning and POI localization highlight fundamental limitations in current MLLM spatial reasoning. We discuss these challenges and outline methodological recommendations for developing more robust LLM‑based indoor navigation systems.



