Wan Juan
收藏资源简介:
万卷(Wan Juan)是一个大规模的多模态中英文数据集,由上海人工智能实验室创建。该数据集包含文本、图文和视频三种模态,总容量超过2TB,其中文本数据超过6亿文档,存储量超过1TB;图文数据处理成文档,总数超过2200万,数据大小超过200GB;视频文件超过1000个,数据大小超过900GB。数据来源于广泛的网络资源,经过算法处理和人工验证确保数据安全、高质量和价值对齐。万卷数据集支持大型模型训练,特别是在多模态任务中,如视频字幕和视频问答,显示出显著优势。
Wan Juan is a large-scale multilingual (Chinese and English) multimodal dataset developed by the Shanghai AI Laboratory. The dataset covers three modalities: text, image-text, and video, with a total capacity exceeding 2 TB. Specifically, the text data includes over 600 million documents occupying more than 1 TB of storage; the image-text data, processed into document format, totals over 22 million with a size exceeding 200 GB; and the video data consists of more than 1,000 files with a total size over 900 GB. The dataset is sourced from a wide range of web resources, and undergoes algorithmic processing and manual verification to ensure data security, high quality, and value alignment. Wan Juan supports the training of large-scale models, and exhibits notable advantages particularly in multimodal tasks such as video captioning and video question answering.

- 1WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models上海人工智能实验室 · 2023年



