遇见数据集

lccshunli/LongVideoBench

收藏
Hugging Face2026-04-26 更新2026-05-03 收录
官方服务:

资源简介:

LongVideoBench是一个用于评估大型多模态模型(LMMs)在长视频理解方面能力的问答基准数据集。它包含3,763个从网络收集的视频,带有字幕,涵盖多样化的主题,视频长度可达一小时。数据集旨在解决从长输入中准确检索和推理详细信息的挑战,并引入了一个称为“引用推理”的新任务,其中问题包含引用相关视频上下文的查询,要求模型对这些细节进行推理。LongVideoBench包括6,678个人工标注的多选题,分布在17个类别中,使其成为长形式视频理解最全面的基准之一。评估显示,即使是先进的专有模型(如GPT-4o、Gemini-1.5-Pro、GPT-4-Turbo)也面临显著挑战,开源模型表现更差。仅当模型处理更多帧时,性能才会提高,这确立了LongVideoBench作为未来长上下文LMMs的重要基准。数据集由LongVideoBench团队策划,语言为英语,许可证为CC-BY-NC-SA 4.0,仅限非商业使用。

LongVideoBench is a question-answering benchmark designed to evaluate large multimodal models (LMMs) on long video understanding. It comprises 3,763 web-collected videos with subtitles across diverse themes, with video lengths up to an hour. The dataset targets the challenge of accurately retrieving and reasoning over detailed information from lengthy inputs, and introduces a novel task called referring reasoning, where questions contain a referring query that references related video contexts, requiring the model to reason over these details. LongVideoBench includes 6,678 human-annotated multiple-choice questions across 17 categories, making it one of the most comprehensive benchmarks for long-form video understanding. Evaluations show significant challenges even for advanced proprietary models (e.g., GPT-4o, Gemini-1.5-Pro, GPT-4-Turbo), with open-source models performing worse. Performance improves only when models process more frames, establishing LongVideoBench as a valuable benchmark for future long-context LMMs. The dataset is curated by the LongVideoBench Team, in English, under the CC-BY-NC-SA 4.0 license, for non-commercial use only.

提供机构:
lccshunli
二维码
社区交流群
二维码
科研交流群
商业服务