OmniCorpus
收藏资源简介:
OmniCorpus是由上海人工智能实验室等机构创建的统一多模态语料库,包含86亿图像和1696亿文本标记,数据规模巨大。该数据集通过高效的数据引擎从英语和非英语网站以及视频平台等多种来源收集,支持多种数据格式,如纯文本、图像-文本对和交错格式。创建过程中,采用了先进的数据处理技术和人工反馈过滤,确保数据质量。OmniCorpus主要应用于多模态大型语言模型的研究,旨在提升模型的理解和生成能力。
OmniCorpus is a unified multimodal corpus created by institutions including Shanghai AI Laboratory and other organizations. It contains 8.6 billion images and 169.6 billion text tokens, with an exceptionally large data scale. Collected from diverse sources such as English and non-English websites, video platforms and more via an efficient data engine, this corpus supports multiple data formats including plain text, image-text pairs and interleaved formats. During its creation, advanced data processing technologies and human feedback-based filtering are adopted to ensure data quality. OmniCorpus is mainly applied to multimodal large language model research, aiming to enhance the models' understanding and generation capabilities.

- 1OmniCorpus: An Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text上海人工智能实验室 · 2024年



