Vision-and-Language Parking Dataset
收藏资源简介:
Autonomous parking represents the final stage in autonomous driving applications. However, current Autonomous Valet Parking (AVP) technology still relies on human-in-the-loop decision-making for parking spots due to the limitations in scene understanding and challenges in multimodal data fusion. This paper defines autonomous parking decision-making as a Vision-and-Language Navigation (VLN) problem, where an autonomous vehicle makes decisions based on visual perception and user command interpretation in parking lots. By formalizing the static object features of typical parking lots, including locations, obstacles and attributes, this paper introduces the Vision-and-Language Parking (VLP) dataset, featuring 174 onboard panoramic images and 11310 structured information from real parking environments at six different time points and two structures, with more than 300 instructions, marking it as the first vision navigation dataset driven by user natural language instructions.



