ucsbai/SAW-Bench
收藏资源简介:
SAW-Bench(真实世界中的情境感知基准测试)是一个用于评估多模态基础模型中“观察者中心”情境感知能力的基准测试,即从自我中心视角随时间推移推理空间、运动和可能行动的能力。与先前强调“环境中心”关系(场景中物体如何相互关联)的基准测试不同,SAW-Bench探究模型是否能从自我中心视频中保持一致的观察者中心空间状态。它包含786个使用Ray-Ban Meta(第二代)智能眼镜捕获的真实世界视频,以及2,071个人工标注的问答对,涵盖六项任务:定位(在空间中的位置)、方向(目标相对于当前朝向的位置)、形状(行进路径的形状)、反向规划(如何返回起点)、记忆(两次访问间场景的变化)和可操作性(从当前姿态可行的行动)。即使最佳模型也落后人类表现37.66%。数据集适用于非商业学术研究,遵循CC BY-NC 4.0许可证。
SAW-Bench (Situated Awareness in the Real World) is a benchmark for evaluating observer-centric situated awareness in multimodal foundation models — the ability to reason about space, motion, and possible actions relative to ones own egocentric viewpoint as it evolves over time. Unlike prior benchmarks that emphasize environment-centric relations (how objects relate to each other in a scene), SAW-Bench probes whether a model can maintain a coherent observer-centric spatial state from egocentric video. It comprises 786 real-world videos captured with Ray-Ban Meta (Gen 2) smart glasses and 2,071 human-annotated question–answer pairs across six tasks: localization, direction, shape, revplan, memory, and affordance. Even the best model trails humans by 37.66%. The dataset is for non-commercial academic research only, under CC BY-NC 4.0 license.




