Internally Curated Dataset
收藏arXiv2025-09-30 收录
数据链接:
官方服务:
资源简介:
该数据集包含了超过10亿张图像-文本对,用于训练D-JEPA⋅T2I模型。在整理过程中,我们排除了美学评分低于5.0的图像,并使用OCR工具过滤掉包含文本的图片。此外,为确保数据的真实性,合成数据集仅占整个数据集的大约5%。该数据集的规模超过10亿图像-文本对,主要用于高分辨率图像合成和文本到图像的生成任务。
This dataset contains over 1 billion image-text pairs for training the D-JEPA⋅T2I model. During the dataset curation stage, we excluded images with an aesthetic score below 5.0 and filtered out images containing text using OCR tools. Additionally, to ensure data authenticity, synthetic datasets only account for approximately 5% of the entire dataset. With over 1 billion image-text pairs in total, this dataset is primarily used for high-resolution image synthesis and text-to-image generation tasks.
提供机构:
Internally Curated搜集汇总
数据集介绍

背景与挑战
背景概述
该数据集与D-JEPA·T2I模型相关,专注于高分辨率文本到图像生成,通过next-token预测和flow matching loss实现连续分辨率学习,但详情页面未提供数据集本身的具体信息(如数据量、来源或格式)。
以上内容由遇见数据集搜集并总结生成



