VAST-27M
收藏资源简介:
VAST-27M是由中国科学院自动化研究所创建的大规模全模态视频字幕数据集,包含2700万个视频片段。每个视频片段都配有视觉、音频和字幕模态的字幕,这些字幕是通过训练单独的视觉和音频字幕生成器以及使用大型语言模型Vicuna-13b整合生成的。数据集的创建旨在推动多模态视频-文本预训练模型的研究,特别是在视觉-文本、音频-文本和全模态视频-文本任务中。VAST-27M数据集的应用领域广泛,包括视频内容理解、视频检索、视频字幕生成和视频问答等。
VAST-27M is a large-scale full-modal video caption dataset developed by the Institute of Automation, Chinese Academy of Sciences, containing 27 million video clips. Each video clip is paired with captions across visual, audio and text modalities, which are generated by training standalone visual and audio caption generators and integrating their outputs via the large language model Vicuna-13b. This dataset is constructed to advance research on multimodal video-text pre-trained models, especially for visual-text, audio-text and full-modal video-text tasks. VAST-27M has a wide spectrum of applications, including video content understanding, video retrieval, video caption generation and video question answering.




