Alibaba-NLP/xvbench
收藏资源简介:
XVBench是一个用于评估多模态检索增强生成系统在跨视频理解任务上的基准数据集,由论文《VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph》引入。该数据集基于HowTo100M视频库构建,HowTo100M是一个大规模叙述性教学视频集合。数据集专注于需要模型或代理从大型视频库中检索和推理视频证据的问题,与单视频问答相比,旨在测试系统是否能找到相关视觉片段、保留细粒度视觉细节,并回答依赖于跨视频片段分布信息的问题。每个数据示例包含一个问题、一个真实答案、源视频标识符以及一个或多个支持答案的参考视频片段。任务为开放式问答,语言为英语,许可证为CC BY 4.0。
XVBench is a benchmark dataset for evaluating multimodal retrieval-augmented generation systems on cross-video understanding tasks, introduced by the paper *VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph*. Built upon the HowTo100M video repository—a large-scale collection of narrated instructional videos—this dataset focuses on scenarios where models or AI agents need to retrieve and reason over video evidence from large-scale video libraries. Compared to single-video question answering, it aims to test whether a system can locate relevant visual segments, retain fine-grained visual details, and answer questions that depend on distributed information across multiple video clips. Each data instance includes a question, a ground-truth answer, source video identifiers, and one or more reference video segments that support the answer. The task is open-ended question answering, with all textual content in English, and the dataset is licensed under CC BY 4.0.



