遇见数据集

Voxel51/NaviTrace

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

NaviTrace是一个视觉问答基准数据集,用于评估视觉语言模型在真实场景中具身导航理解的能力。每个样本包含一张第一人称视角的户外环境图像和一条自然语言导航指令,任务是根据给定的具身类型(人类、腿式机器人、轮式机器人或自行车)预测一个二维导航轨迹(图像空间中的一系列路径点)。数据集包含1002个场景,提供超过3000条专家标注的轨迹,涵盖多种导航挑战类别,如几何地形、语义地形、可访问性、可见性、社会规范、动态障碍物避免和静态障碍物避免。每个样本还包括场景元数据,如地理位置、环境类型、光照和天气条件。数据集主要用于基准测试、评估空间基础和目标定位,以及研究具身类型对导航决策的影响。

NaviTrace is a Visual Question Answering benchmark for evaluating how well vision-language models (VLMs) understand embodiment-specific navigation in real-world scenes. Each sample presents a first-person image of an outdoor environment paired with a natural language navigation instruction. The task is to predict a 2D navigation trace — a sequence of waypoints in image space — that a given embodiment type would follow to complete the instruction. The dataset contains 1,002 scenarios with over 3,000 expert-annotated traces across four embodiment types: Human, Legged Robot, Wheeled Robot, and Bicycle. It includes various navigation challenge categories, such as Geometric Terrain, Semantic Terrain, Accessibility, Visibility, Social Norms, Dynamic Obstacle Avoidance, and Stationary Obstacle Avoidance. Each sample also contains scene-level metadata covering geographic location, environment type, lighting, and weather. The dataset is intended for benchmarking VLMs on embodied navigation trace prediction, evaluating spatial grounding and goal localization, and studying how embodiment type affects navigation decisions.

提供机构:
Voxel51
二维码
社区交流群
二维码
科研交流群
商业服务