SocialNav-SUB
收藏资源简介:
SocialNav-SUB是一个视觉问答(VQA)数据集和基准测试,旨在评估视觉语言模型(VLM)在现实世界中社交机器人导航场景下的场景理解能力。数据集包含4968个独特的问题及其对应的答案,这些答案由人类提供,作为真实标签。SocialNav-SUB基于SCAND数据集构建,提供了各种人群密度和社会导航交互的社会机器人导航场景,并具有丰富的对象中心视觉表示,包括机器人的视觉视角和包含行人坐标跟踪的鸟瞰图(BEV)。该数据集旨在解决社会机器人导航中的空间推理、时序推理和理解复杂人类意图等关键挑战。
SocialNav-SUB is a visual question answering (VQA) dataset and benchmark designed to evaluate the scene understanding capabilities of vision-language models (VLMs) in real-world social robot navigation scenarios. The dataset contains 4,968 unique questions paired with their human-provided ground-truth answers. Built upon the SCAND dataset, SocialNav-SUB provides social robot navigation scenarios with varying crowd densities and social navigation interactions, and features rich object-centric visual representations, including the robot's egocentric view and a bird's-eye view (BEV) with tracked pedestrian coordinates. This dataset aims to address key challenges in social robot navigation, such as spatial reasoning, temporal reasoning, and understanding complex human intentions.




