VideoMind
收藏资源简介:
VideoMind是一个包含103K视频样本的多模态视频数据集,每个样本都伴有音频和详尽的文本描述。数据集的内容包括图像、视频、音频和文本,以及自动语音识别(ASR)、光学字符识别(OCR)和各种语义标签。VideoMind旨在通过提供全面且深入的视频内容文本解释,促进视频理解和增强多模态表示。数据集的创建过程包括从社交媒体平台选择视频,并利用mLLM生成从事实到意图的多层次文本描述。VideoMind适用于需要深入理解视频内容的领域,如情绪和意图识别,并通过混合认知检索实验评估模型的视频理解能力。
VideoMind is a multimodal video dataset containing 103K video samples, each paired with corresponding audio and exhaustive textual descriptions. The dataset encompasses images, videos, audios, texts, as well as automatic speech recognition (ASR), optical character recognition (OCR) and various semantic labels. VideoMind seeks to advance video understanding and boost multimodal representation learning by providing comprehensive and in-depth textual explanations of video content. The dataset is developed by selecting source videos from social media platforms, and generating multi-level textual descriptions ranging from factual content to intentions via mLLMs. VideoMind is applicable to domains requiring in-depth comprehension of video content, such as emotion and intention recognition, and can be utilized to evaluate models' video understanding capabilities through mixed cognitive retrieval experiments.
VideoMind数据集概述
数据集简介
- 名称:VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
- 论文地址:https://arxiv.org/abs/2507.18552
- 版本:V1版本(包含视频注释和黄金标准基准)
数据集内容
- 样本数量:103K视频样本(其中3K仅用于测试)
- 数据类型:
- 每个视频样本包含音频数据
- 系统且详细的文本描述(三个层次:事实层、抽象层和意图层)
- 文本总量:超过2200万词,平均每个样本约225词
数据特点
- 独特特征:提供意图表达,需通过整合整个视频的上下文进行推测
- 标注内容:包括主题、地点、时间、事件、动作和意图等
- 黄金标准基准:包含3000个经过人工验证的样本
数据统计
- 视频统计:包含在数据集中(具体统计信息见原始图片)
- 上传者意图词云:展示上传者意图的词汇分布
- 角色意图词云:展示角色意图的词汇分布
下载信息
- 视频注释下载:
- OpenDataLab:https://opendatalab.com/Dixin/VideoMind
- HuggingFace:https://huggingface.co/datasets/DixinChen/VideoMind
- 基准视频下载:https://drive.google.com/file/d/1RbEjY1_glJ8yEwn1f5SXGs5kCn6uAqvY/view?usp=drive_link
引用信息
bibtex @misc{yang2025videomindomnimodalvideodataset, title={VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding}, author={Baoyao Yang and Wanyun Li and Dixin Chen and Junxiang Chen and Wenbin Yao and Haifeng Lin}, year={2025}, eprint={2507.18552}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2507.18552}, }




