LongVQUBench
收藏资源简介:
LongVQUBench是由南洋理工大学创建的一个综合性基准数据集,专门用于评估大规模视觉语言模型在长期视频质量理解方面的能力。该数据集包含超过1200个视频,涵盖电影、纪录片、监控录像、第一人称记录和动画等多种类型,视频时长从几分钟到近两小时不等,并配备了1500个多项选择和开放式问答对,以支持验证和测试。其构建过程基于从公开数据集和媒体库中收集的长视频,通过系统性地应用空间和时间失真来模拟真实退化模式,并采用分层评估框架和针孔失真问答范式来探究模型的感知推理能力。该数据集主要应用于视频质量评估和人工智能领域,旨在解决现有基准在长视频质量理解方面的不足,推动模型在感知保真度、时间连贯性和累积退化方面的进步,以实现更接近人类水平的长期视频理解。
LongVQUBench is a comprehensive benchmark dataset created by Nanyang Technological University, specifically designed to evaluate the capabilities of large-scale vision-language models in long-form video quality understanding. This dataset contains over 1200 videos spanning diverse genres including films, documentaries, surveillance footage, first-person recordings, animations and more, with durations ranging from several minutes to nearly two hours. It is equipped with 1500 multiple-choice and open-ended question-answer pairs to support model validation and testing. The dataset is built upon long videos collected from public datasets and media libraries, where systematic spatial and temporal distortions are applied to simulate real-world degradation patterns. A hierarchical evaluation framework and the pinhole distortion-based question-answering paradigm are adopted to explore the perceptual reasoning abilities of models. This dataset is mainly applied in the fields of video quality assessment and artificial intelligence, aiming to address the shortcomings of existing benchmarks in long-form video quality understanding, and promote the advancement of models in terms of perceptual fidelity, temporal coherence and cumulative degradation, so as to achieve long-form video understanding that is closer to human-level performance.
LongVQUBench 数据集详情
LongVQUBench 是一个用于评估大视觉语言模型(LVLMs)长时视频质量理解能力的基准测试。该论文已被 ECCV 2026 收录。
核心信息
- 数据规模:包含 1,200 个多样化视频,时长平均约 12分钟;以及 1,500 个用于验证和测试的问答对(包括多项选择题和开放性问题)。
- 视频来源:涵盖电影、纪录片、监控录像、自我中心视角录像、动画、信息图、新闻、vlog 和烹饪等多种类型。部分视频源自 LongVideoBench、MLVU 和 LongVideo-Reason-Eval 等数据集。
评估框架
数据集设计了三个层级递进的评估级别,用以测试模型从局部失真检测到全局感知推理的能力:
- 局部事件质量理解(LQU):检测、定位并评估单一、有界时域内的失真事件(如局部模糊、闪烁或压缩噪声)的严重程度。
- 跨事件质量推理(CQR):比较、关联并整合分布在较长时域跨度内的多个失真事件,评估其累积效应和时间关系。
- 全局质量理解(GQU):对整个视频进行全局感知判断,跟踪质量趋势,识别主要退化类型,并评估整体感知稳定性。
此外,一个名为 针式失真问答(NDQA) 的任务范式被嵌入到上述所有三个层级中,通过在视频中稀疏插入微弱的空间或时间伪影来探测模型的细粒度检测和推理能力。
实验评估
对 14 个最先进的LVLMs进行了广泛实验,结果显示:随着视频长度和推理深度的增加,模型性能显著下降,表明它们在长程时间整合和感知归因方面能力有限。
相关资源
- 数据集:🤗 数据集(根据页面信息推断)
- 代码:GitHub(根据页面信息推断)
- 排行榜:Leaderboard(根据页面信息推断)




