jasperai/monet
收藏资源简介:
MONET(大规模、开放、非冗余和丰富的文本到图像数据集)是一个专为训练文本到图像(T2I)系统而设计的大规模、经过筛选的图像-文本数据集。它包含从九个异构开放源(6个真实源和3个合成源)的29亿原始对中通过多阶段安全过滤、基于域的过滤、精确和近似重复去除以及使用多个视觉语言模型重新标注而得到的1.049亿高质量图像-文本对,并进一步通过合成生成的样本进行增强。每张图像都附有预计算的嵌入、结构化注释和预编码的VAE潜在表示,以加速下游使用。
MONET (Massive, Open, Non-redundant, and Enriched Text-to-Image Dataset) is a large-scale curated image-text dataset specifically designed for training text-to-image (T2I) systems. It comprises 104.9 million high-quality image-text pairs derived from 2.9 billion raw pairs across nine heterogeneous open sources (6 real-world sources and 3 synthetic sources) via multi-stage safety filtering, domain-based filtering, exact and approximate deduplication, and re-annotation using multiple vision-language models. The dataset is further augmented with synthetically generated samples. Each image is accompanied by pre-computed embeddings, structured annotations, and pre-encoded VAE latent representations to accelerate downstream usage.




