data-archetype/cc12_imagenet21k_512_subset
收藏资源简介:
`cc12_imagenet21k_512_subset`是一个基于CC12/ImageNet21K的512-base分桶子集数据集,主要用于文本到图像的训练或数据集检查。数据集包含1,655,489张图片,这些图片满足以下条件:拥有OK的`caption_gemini`标题、JPEG格式的源文件适合JPEG直通导出,并且足够大以适应512-base分桶而无需放大。数据集采用`bucketed_shards_v1`格式,每个样本包含三个文件:JPEG图像、UTF-8标题文本和每样本元数据。数据集的分桶策略基于SDXL风格的长宽比原型桶,基础分辨率为512。
`cc12_imagenet21k_512_subset` is a 512-base bucketed export of a larger CC12/ImageNet21K recap dataset, intended for text-to-image training or dataset inspection with WebDataset-style loaders. It contains 1,655,489 images that meet the following criteria: have an OK `caption_gemini` caption, are backed by JPEG-family source files suitable for JPEG passthrough export, and are large enough to fit a 512-family bucket without upscaling. The dataset uses the `bucketed_shards_v1` format, with each sample stored as three files: JPEG image bytes, UTF-8 caption text, and per-sample metadata. The bucket family is SDXL-style aspect-ratio proto buckets defined at 512 base.




