VIABench
收藏资源简介:
VIABench是一个全面的以自我为中心视角的视频基准数据集,旨在评估多模态大语言模型在现实世界视觉障碍辅助场景中的性能。数据集收集自盲人录制或分享的真实视频,包含761个视频,总时长46.9小时,并提供了14,526个手动精心标注的样本。它围绕三个核心任务构建:主动提醒(要求模型理解实时视频流,预测对导航关键的事件,并在事件发生前提供及时的语言提醒)、视觉问答(要求模型回答用户提出的关于视频中环境或物体的问题)以及视觉引导交互(评估模型在用户与环境之间完成有意图的交互时,进行上下文感知推理和指导的能力)。这些任务覆盖了视觉障碍辅助中的关键需求,适用于视频到文本生成、场景理解、实时决策支持等研究方向。数据集采用MIT许可协议,主要任务类别为视频-文本到文本。
VIABench is a comprehensive egocentric video benchmark dataset designed to evaluate the performance of multimodal large language models in real-world visual impairment assistance scenarios. The dataset is collected from real videos recorded or shared by blind individuals, containing 761 videos with a total duration of 46.9 hours, and provides 14,526 manually and meticulously annotated samples. It is built around three core tasks: proactive reminding (requiring the model to understand real-time video streams, predict events critical to navigation, and provide timely verbal reminders before events occur), visual question answering (requiring the model to answer user-posed questions about the environment or objects in the video), and visual-guided interaction (evaluating the models ability to perform context-aware reasoning and guidance when users complete intentional interactions with the environment). These tasks cover key needs in visual impairment assistance and are applicable to research directions such as video-to-text generation, scene understanding, and real-time decision support. The dataset is released under the MIT license, with the primary task category being video-text-to-text.
VIABench 数据集概述
基本信息
- 许可证:MIT
- 任务类别:视频-文本到文本
- 所属机构:MCG-NJU
数据集简介
VIABench 是一个面向视觉障碍辅助场景的全面第一人称视频基准数据集,用于评估多模态大语言模型在实际环境中的表现。数据来源于盲人拍摄或分享的视频。
数据集规模
- 视频数量:761 个
- 总时长:46.9 小时
- 标注数量:14,526 条人工精心策划的标注
核心任务
数据集围绕以下三个核心任务构建:
-
主动提醒(Proactive Reminder)
- 评估模型是否能够理解持续的视频流,预测导航中的关键事件,并在事件发生前及时提供口头提醒。
-
视觉问答(Visual Question Answering, VQA)
- 评估模型是否能够回答用户提出的关于视频环境或物体的问题。
-
视觉引导交互(Vision-Guided Interaction)
- 评估模型是否能够进行上下文感知的推理和指导,帮助用户完成与环境的意图性交互。
相关资源
- 论文:https://arxiv.org/abs/2607.14660
- 代码仓库:https://github.com/MCG-NJU/VIABench




