MINT-1T
收藏资源简介:
MINT-1T是一个开源的多模态交错数据集,包含一万亿文本token和三十亿张图像,是现有开源数据集规模的10倍。此外,还包括了如PDF和ArXiv论文等之前未被充分利用的资源。目前正在进行最后的完善工作,并即将开源该数据集。
MINT-1T is an open-source multimodal interleaved dataset, encompassing one trillion text tokens and three billion images, making it ten times the scale of existing open-source datasets. Additionally, it includes previously underutilized resources such as PDFs and ArXiv papers. The dataset is currently undergoing final refinements and is set to be open-sourced soon.
MINT-1T数据集概述
数据集简介
MINT-1T是一个开放源代码的多模态交错数据集,包含一万亿文本令牌和三十亿图像,相比现有开放源代码数据集规模扩大了10倍。此外,该数据集还包含了先前未充分利用的资源,如PDF文件和ArXiv论文。目前,MINT-1T数据集的最终调整工作正在进行中,预计不久将开放源代码。
数据集特点
- 规模:包含一万亿文本令牌和三十亿图像。
- 多模态:支持文本与图像的交错。
- 资源:包含PDF文件和ArXiv论文等未充分利用的资源。
更新信息
已发布技术报告,详细信息可参考技术报告。
引用信息
若您发现本数据集对您的工作有用,请考虑引用以下文献:
@article{awadalla2024mint1t, title={MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens}, author={Anas Awadalla and Le Xue and Oscar Lo and Manli Shu and Hannah Lee and Etash Kumar Guha and Matt Jordan and Sheng Shen and Mohamed Awadalla and Silvio Savarese and Caiming Xiong and Ran Xu and Yejin Choi and Ludwig Schmidt}, year={2024} }




