CG-AV-Counting
收藏资源简介:
CG-AV-Counting是一个针对多模态大型语言模型(MLLMs)的长视频计数能力评估的基准数据集。该数据集由南京大学的研究团队创建,包含1,027个多模态问题,5,845个注解线索,涵盖497个长视频。数据集支持黑盒和白盒评估,用于全面测试端到端和基于推理的计数能力。CG-AV-Counting旨在解决现有计数基准在视频时长、查询类型、线索注解和模态覆盖方面的限制。
CG-AV-Counting is a benchmark dataset for evaluating the long-video counting capabilities of multimodal large language models (MLLMs). This dataset was created by a research team from Nanjing University, containing 1,027 multimodal questions and 5,845 annotated cues, covering 497 long videos. The dataset supports both black-box and white-box evaluations, enabling comprehensive testing of end-to-end and reasoning-based counting capabilities. CG-AV-Counting aims to address the limitations of existing counting benchmarks in terms of video duration, query types, cue annotation, and modal coverage.
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
数据集概述
- 数据集名称: CG-AV-Counting
- 作者: Lidong Lu, Guo Chen, Zhiqi Li, Yicheng Liu, Tong Lu
- 机构: 南京大学
- 年份: 2025
- 论文链接: https://arxiv.org/abs/2506.05328
数据集详情
- 视频数量: 497个长视频(均超过10分钟)
- 问题数量: 1,027个多模态查询问题
- 标注线索: 5,845个细粒度人工标注线索
- 模态需求: 40%样本需音频+视觉模态,其余仅需视觉模态
- 计数目标类型: 对象计数、事件计数、属性计数
- 数值范围: 1-76(长尾分布,主要集中在1-20)
- 视频类别: 超过10类(体育、生活记录、幽默、教程等)
基准测试特点
- 首创性: 首个专门评估MLLMs视频计数能力的综合基准
- 对比优势:
- 同时包含音频和视觉模态
- 更复杂的查询
- 提供细粒度计数线索
- 支持端到端和基于推理的计数评估
评估指标
黑盒评估(端到端计数)
- Long Acc: 完整视频中的计数和时序定位
- Ref Acc: 修剪参考片段中的计数
- 指标:
- 准确率(Acc)
- 容错准确率(OBOA)
- 平均绝对误差(MAE)
- 均方根误差(RMSE)
白盒评估(显式证据定位)
- White-box Counting Score (WCS): 结合定位准确性和计数惩罚
- Instruction-Following Accuracy (IFA): 输出格式匹配比例
基线模型AV-Reasoner
- 训练方法: GRPO+课程学习
- 性能: 在多个基准测试中达到SOTA
- 局限性: 在域外基准上,语言空间推理未能带来性能提升
引用格式
bibtex @misc{lu2025avreasoner, title={AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs}, author={Lidong Lu and Guo Chen and Zhiqi Li and Yicheng Liu and Tong Lu}, year={2025}, eprint={2506.05328}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2506.05328}, }




