VidText
收藏资源简介:
VidText是一个专为多模态大语言模型(MLLMs)在视频文本理解方面的系统评估而设计的综合基准。它涵盖了27个细粒度类别的多样化视频,包括多种语言和场景,并设计了8个任务,涵盖感知和推理维度。这些任务挑战MLLMs在不同粒度上利用视频中动态出现的文本线索——从整体视频级理解到实例级定位。
VidText is a comprehensive benchmark designed for the systematic evaluation of Multimodal Large Language Models (MLLMs) in video-text understanding. It encompasses diverse videos spanning 27 fine-grained categories across multiple languages and scenarios, and incorporates 8 tasks covering both perceptual and reasoning dimensions. These tasks challenge MLLMs to leverage dynamically emerging textual cues in videos at varying granularities, ranging from holistic video-level comprehension to instance-level localization.
VidText数据集概述
数据集简介
VidText是一个专为多模态大语言模型(MLLMs)视频文本理解能力评估设计的综合基准测试集,涵盖27个细粒度类别的多样化视频内容,支持多语言和多场景评估。
关键特性
- 数据多样性:包含不同长度的视频,覆盖27个细粒度类别
- 多语言支持:涵盖多种语言和场景
- 任务设计:包含8个任务,覆盖感知和推理两个维度
- 评估粒度:从视频级整体理解到实例级定位
技术指标
- 模型表现:当前最佳模型平均准确率仅为45.3%
- 评估维度:
- 整体OCR识别
- 整体推理
- 局部OCR识别
- 局部推理
- 文本定位
- 时序因果推理
- 文本追踪
- 空间推理
数据获取
- 标注文件:https://github.com/Naxyang/VidText/tree/master/data
- 原始视频:https://huggingface.co/datasets/sy1998/VidText
使用许可
- 许可证类型:CC-BY-NC-SA-4.0
- 使用限制:仅限研究用途,禁止商业用途
引用信息
bibtex @article{VidText, title={VidText: Towards Comprehensive Evaluation for Video-Text Understanding}, author={Yang, Zhoufaran and Shu, Yan and Yang, Zhifei and Zhang, Yan and Li, Yu and Lu, Keyang and Zeng, Gangyan and Liu, Shaohui and Zhou, Yu and Sebe, Nicu}, journal={arXiv preprint arXiv:2505.22810}, year={2025} }
评估资源
- 评估详情请参考:https://github.com/shuyansy/VidText/data/evaluation




