MINT-1T
收藏资源简介:
MINT-1T是由华盛顿大学和Salesforce Research合作创建的开放源代码多模态交错数据集,包含一万亿文本令牌和三十亿图像,是目前最大且最多样化的开放源代码多模态交错数据集。数据集内容丰富,涵盖HTML、PDF和ArXiv等多种来源,旨在通过提供大规模、多样化的训练数据,推动前沿大型多模态模型(LMMs)的发展,解决现有开放源代码多模态数据集规模和多样性不足的问题。
MINT-1T is an open-source multimodal interleaved dataset co-developed by the University of Washington and Salesforce Research. It contains 1 trillion text tokens and 3 billion images, making it the largest and most diverse open-source multimodal interleaved dataset to date. The dataset features rich content sourced from various formats and platforms including HTML, PDF, ArXiv, and more. It aims to advance the development of cutting-edge large multimodal models (LMMs) by providing large-scale and diverse training data, addressing the shortcomings of existing open-source multimodal datasets in terms of scale and diversity.

- 1MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens华盛顿大学 · 2024年



