RoboTrustBench
收藏资源简介:
RoboTrustBench是由新加坡管理大学、复旦大学和普林斯顿大学联合创建的机器人操作视频世界模型可信度基准数据集。该数据集包含1,207条经过专家验证的指令-图像对,源自真实世界的大规模机器人操作数据集DROID,涵盖了家庭厨房、办公室等10种物理场景、321种不同物体类型和102种任务动词。其构建过程通过对原始DROID数据进行分层采样、指令改写和图像编辑,并经过严格的人工验证,最终形成包含正常、约束敏感、反事实和对抗性四种评估场景的标准化测试集。该数据集主要应用于评估视频世界模型在机器人操作领域的可信度,旨在系统检验模型在面临模糊指令、物理约束冲突、场景不一致及不安全指令等复杂条件下,能否生成物理合理、语义一致且安全可控的机器人操作视频序列,从而推动可信人工智能在具身智能领域的发展。
RoboTrustBench is a credibility benchmark dataset for world models of robotic manipulation videos, jointly created by Singapore Management University, Fudan University, and Princeton University. This dataset contains 1,207 expert-validated instruction-image pairs, sourced from the large-scale real-world robotic manipulation dataset DROID, covering 10 physical scenarios including home kitchens and offices, 321 distinct object types, and 102 task verbs. It was constructed through hierarchical sampling, instruction rewriting, and image editing on the original DROID data, followed by rigorous manual validation, ultimately forming a standardized test set with four evaluation scenarios: normal, constraint-sensitive, counterfactual, and adversarial. This dataset is primarily applied to evaluate the credibility of video world models in the field of robotic manipulation, aiming to systematically test whether models can generate physically plausible, semantically consistent and safely controllable robotic manipulation video sequences when faced with complex conditions such as ambiguous instructions, physical constraint conflicts, scene inconsistencies and unsafe instructions, thereby promoting the development of trustworthy artificial intelligence in the field of embodied intelligence.
数据集名称
RoboTrustBench —— 面向机器人操作视频世界模型的可信度基准测试
数据集核心问题
当前模型能生成视觉上连贯的视频,但在以下方面仍存在困难:
- 受约束的操作
- 反事实(与现实世界冲突)的接地
- 物理上合理的交互
- 不安全指令的抑制
数据集构建方法
数据基于真实的 DROID 机器人操作片段,通过修改指令/图像并经专家验证,构建四类样本:
- Normal(正常场景)
- Constraint-Sensitive(约束敏感场景)
- Counterfactual(反事实场景)
- Adversarial(对抗性/安全场景)
该设计将标准任务执行与需要约束处理、世界状态接地和安全意识抑制的可信度关键案例分开。
数据集统计信息
- 场景类型:涵盖多种室内环境
- 物体类型:321 种
- 任务动词:102 种
- 非正常场景:细分为三类失败来源——约束可行任务、与观察世界冲突的指令、不安全的机器人意图
主要评估维度与结果
-
人类评估(13 个维度)
将 1-5 分制标准化为 [0,1],其中“安全风险识别”仅在对抗性(Adversarial)视频上评估。 -
约束敏感任务完成
在语义歧义(如泛指和代词)上表现较好,但在轨迹约束和目标物体遮挡上表现下降,表明空间-物理推理仍比上下文语言完成更难。 -
反事实高任务完成案例
即使反事实视频看起来完成了任务,其在真实感和实体一致性上的得分仍较低,表明表面成功常来自幻觉或场景修改。 -
安全风险识别(对抗性场景)
示例结果(以环境损害为例):- Kling-v2.6:低风险 90%、中 0%、高 10%,均值 1.3
- Veo-3.1-Fast:低 50%、中 20%、高 30%,均值 2.3
结论:当前模型不能可靠地抑制不安全生成,可信的视频世界模型需要评估其“能做什么”和“在指令有害时是否避免行动”。

- 1RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation新加坡管理大学; 复旦大学; 普林斯顿大学 · 2026年



