NavTrust
收藏资源简介:
NavTrust是由加州大学河滨分校等机构联合提出的首个统一基准,旨在系统评估VLN和OGN代理在现实损坏条件下的可信度。该数据集基于Habitat-Matterport3D、R2R和RxR数据集构建,包含八类RGB图像损坏(如运动模糊、低光照)、四类深度传感器损坏(如高斯噪声、多路径效应)以及五类指令损坏(如风格变异、恶意提示)。通过对比清洁与损坏场景下的性能差异,该数据集揭示了现有SOTA导航模型的脆弱性,并为开发更鲁棒的具身智能系统提供了标准化测试平台。其创新性在于首次统一了多模态损坏评估框架,重点关注了传统研究忽视的深度传感器退化问题。
NavTrust is the first unified benchmark jointly proposed by the University of California, Riverside and other institutions, which aims to systematically evaluate the trustworthiness of Vision-and-Language Navigation (VLN) and Object-Goal Navigation (OGN) agents under realistic corrupted conditions. This dataset is constructed based on Habitat-Matterport3D, R2R, and RxR datasets, and includes eight types of RGB image corruptions (e.g., motion blur, low-light conditions), four types of depth sensor corruptions (e.g., Gaussian noise, multipath effect), and five types of instruction corruptions (e.g., style variation, malicious prompts). By comparing the performance differences between clean and corrupted scenarios, this dataset reveals the vulnerability of current state-of-the-art (SOTA) navigation models, and provides a standardized testbed for developing more robust embodied intelligent systems. Its innovation lies in establishing the first unified multimodal corruption evaluation framework, focusing on the depth sensor degradation issue that has been overlooked by traditional research.
NavTrust: Benchmarking Trustworthiness for Embodied Navigation
数据集概述
NavTrust是一个统一的基准测试,旨在系统性地评估具身导航模型的信任度。它通过在现实场景中破坏输入模态(包括RGB、深度和指令),并评估其对导航性能的影响,以揭示现有模型在真实世界条件下的鲁棒性差距。
关键信息
- 研究领域:具身导航,主要包括视觉语言导航(VLN)和面向目标的导航(OGN)。
- 核心问题:现有工作主要在理想条件下评估模型性能,忽视了现实环境中可能出现的破坏,导致模型在轻微的语言扰动、低光照或运动模糊等情况下表现不可靠。
- 基准测试内容:NavTrust首次在统一框架中,将具身导航智能体暴露于多样化的RGB-Depth破坏和指令变体中。
- 评估方法:对七种最先进的方法进行了广泛评估,揭示了在现实破坏下成功率的大幅下降。
- 缓解策略:系统评估了四种不同的缓解策略以增强鲁棒性:数据增强、师生知识蒸馏、保护性大语言模型(LLM)和轻量级适配器调优。
评估结果
- 破坏类型:包括RGB破坏、深度破坏和指令破坏。
- 性能指标:成功率(SR)和SPL。同时使用基于SR和SPL的性能保持分数(PRS)来量化对破坏的鲁棒性。
- 缓解策略效果:在R2R数据集上测试了数据增强、师生蒸馏和适配器等策略对不同破坏(如低光照)下成功率的影响。同时评估了保护性LLM在R2R数据集上针对不同指令变体的成功率。
引用信息
- 标题:NavTrust: Benchmarking Trustworthiness for Embodied Navigation
- 作者:Yash Chaudhary, Huaide Jiang, Yuping Wang, Raghav Sharma, Manan Mehta, Lichao Sun, Zhiwen Fan, Zhengzhong Tu, Jiachen Li
- 年份:2026
- 论文链接:https://openreview.net/forum?id=ANbAB0tXv3
- 代码状态:即将发布
相关机构
- 加州大学河滨分校
- 密歇根大学
- Workday
- 南加州大学
- 利哈伊大学
- 德州农工大学
- 可信自主系统实验室(TASL)

- 1NavTrust: Benchmarking Trustworthiness for Embodied Navigation加州大学河滨分校·可信自主系统实验室; 密歇根大学; Workday; 南加州大学; 德州农工大学; 利哈伊大学 · 2026年



