DIVE (Dense Information Video Evaluation) Benchmark
收藏资源简介:
DIVE(密集信息视频评估)是首个专为密集视频理解设计的基准,专注于问答驱动的高帧率理解,其中答案相关信息几乎出现在每一帧中。该基准适用于教育/讲座视频、手术程序、手语等有用内容密集分布在帧间的场景。
DIVE (Dense Information Video Evaluation) is the first benchmark specifically designed for dense video understanding. It focuses on question-answering-driven high-frame-rate understanding, where answer-relevant information appears in almost every frame. This benchmark is applicable to scenarios where useful content is densely distributed across frames, such as educational/lecture videos, surgical procedures, sign language and other similar cases.
Dense Information Video Evaluation (DIVE) 数据集概述
数据集简介
DIVE (Dense Information Video Evaluation) 是首个专注于密集视频理解的基准测试,重点研究QA驱动的高帧率理解任务,其中答案相关信息几乎出现在每一帧中。
核心特点
- 密集视频理解:针对有用内容密集分布在帧间的场景(如教育/讲座视频、外科手术、手语)
- 高帧率要求:需要帧级密集推理的问答任务
- 解决现有问题:现有VLLM流水线为控制令牌成本而进行激进下采样,会丢失关键时间细节
技术方法
Gated Residual Tokenization (GRT)
- 运动门控令牌化:通过运动线索检测静态区域并在令牌化过程中跳过,实现相对于FPS的次线性令牌/时间增长
- 语义场景令牌合并:在场景内合并冗余令牌同时保留动态语义
数据集内容
- 任务类型:密集视频问答 (Dense Video QA)
- 数据分割:目前仅发布测试集
- 访问地址:https://huggingface.co/datasets/haichaozhang/DenseVideoEvaluation
使用方法
数据加载
python from datasets import load_dataset ds = load_dataset("haichaozhang/DenseVideoEvaluation", split="test")
评估集成
- 正在准备将DIVE集成到LMMS-EVAL视觉语言模型测试工具包中
- 支持通过LMMS-EVAL框架进行评估
发布计划
- ✅ 2025/09/18:发布DIVE基准测试(测试集)
- ⭕ 将DIVE合并到LMMS-EVAL中
- ⭕ 发布数据集的多FPS版本
- ⭕ 添加更多密集视频任务类别
- ⭕ 发布完整的GRT模型和训练/推理代码
相关资源
- 论文:https://arxiv.org/pdf/2509.14199
- 项目网站:https://zhanghaichao.xyz/DenseVideoUnderstand/
- 代码仓库:https://github.com/hai-chao-zhang/DenseVideoUnderstand/
许可信息
- 数据集:OpenRAIL许可(详见数据集卡片中的条款)
- 代码:随模型发布时公布
引用信息
bibtex @article{zhang2025dive, title={Dense Video Understanding with Gated Residual Tokenization}, author={Haichao Zhang and Wenhao Chai and Shwai He and Ang Li and Yun Fu}, journal={arXiv preprint arXiv:2509.14199}, year={2025} }




