LONGVIDEOBENCH
收藏资源简介:
LONGVIDEOBENCH是一个针对长时视频-语言交错理解的问题回答基准,包含3,763个不同长度的网络收集视频及其字幕,涵盖多种主题。数据集设计用于全面评估大型多模态模型在长期多模态理解上的能力。数据集包含6,678个人工标注的多项选择题,分为17个细粒度类别,旨在测试模型在长视频中的详细多模态信息检索和推理能力。数据集的应用领域包括电影、新闻、生活和知识,旨在解决长视频内容理解中的复杂问题。
LONGVIDEOBENCH is a question answering benchmark targeting long-form video-language interleaved understanding. It contains 3,763 web-collected videos with varying lengths and their corresponding subtitles, covering a wide range of topics. This benchmark is designed to comprehensively evaluate the long-term multimodal understanding capabilities of large multimodal models. It includes 6,678 manually annotated multiple-choice questions divided into 17 fine-grained categories, which aim to test the model's abilities in detailed multimodal information retrieval and reasoning within long videos. The dataset covers application domains such as film, news, daily life and knowledge, and is intended to address complex challenges in long-form video content understanding.




