VideoChat3-Academic2M, VideoChat3-LV116K, VideoChat3-OL617K
收藏资源简介:
VideoChat3数据集是由南京大学、上海人工智能实验室等机构联合构建的高质量多模态指令数据集,旨在推动通用视频理解模型的发展。该数据集包含VideoChat3-Academic2M、VideoChat3-LV116K和VideoChat3-OL617K三个子集,总计300万条样本,覆盖通用视频、长视频和流式视频场景,数据来源于合成与筛选相结合的大规模多模态管道。其创建过程通过可扩展的数据合成流程实现,确保了多样性和高质量标注。该数据集主要应用于视频多模态大语言模型的训练,旨在解决模型在跨视频类型泛化、计算效率及可复现性方面的挑战,为现实世界视频理解系统提供开源基础。
The VideoChat3 dataset is a high-quality multimodal instruction dataset jointly constructed by Nanjing University, Shanghai AI Laboratory and other institutions, aiming to advance the development of general-purpose video understanding models. It comprises three subsets: VideoChat3-Academic2M, VideoChat3-LV116K and VideoChat3-OL617K, with a total of 3 million samples covering general video, long-form video and streaming video scenarios. The dataset is derived from a large-scale multimodal pipeline that integrates data synthesis and filtering. Its construction is implemented via a scalable data synthesis workflow, ensuring data diversity and high-quality annotations. This dataset is primarily applied to the training of video multimodal large language models (LLMs), with the purpose of addressing the challenges of cross-video-type generalization, computational efficiency and reproducibility, and providing an open-source foundation for real-world video understanding systems.

- 1VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding南京大学; 上海人工智能实验室; 南洋理工大学; 北京大学 · 2026年




