ramanv-image-foundation
收藏资源简介:
ramanv-image-foundation 是一个生产级、高可扩展的数据集工程平台,专门用于处理和构建训练及微调文本到图像生成模型(如FLUX、SDXL以及图像编辑模型)所需的大规模数据集。该平台旨在处理超过1000万张图像,支持多GPU图像描述生成、光学字符识别文本提取、CLIP图文对齐度评分、LAION美学评分、重复检测、质量过滤以及自动上传至Hugging Face Hub等功能。数据集的核心内容组织在86个明确的类别文件夹中,涵盖广泛的主题领域,包括设计与布局(如营销材料、品牌标识、海报、广告、图标、插画、3D内容、包装、社交媒体素材)、商业与产品(如电子产品、时尚、化妆品、食品)、交通工具(如汽车、摩托车、飞机、船只)、空间与建筑(如建筑、房地产、室内设计、办公空间、旅行、酒店)、自然与景观(如风景、森林、山脉、海滩、河流、太空)、科学与技术、文化与人(如节日、寺庙、婚礼、城市街景、人像肖像)以及文档与文本(如收据、表格、书籍、菜单、报纸、标牌、证书、手写内容)。数据以多种形式存储和提供:原始图像文件(.jpg/.png)按类别存放;原始描述文本文件;符合特定模式的元数据JSON文件(包含49个字段);用于训练、验证和测试的数据集清单文件(JSON Lines格式);预计算的CLIP/SigLIP特征向量;生成的质量评估报告和预览缩略图。数据处理流程标准化,包括生成描述、提取OCR、计算对齐度和美学分数、执行去重和质量过滤、以及划分数据集分区。该数据集特别适用于训练需要高质量、多样化视觉内容和丰富文本描述的最新文本到图像模型。针对FLUX模型,建议利用详细的长描述、根据宽高比分桶训练以及应用美学分数过滤(例如,分数低于6.0的图像)。针对SDXL模型,建议使用CLIP对齐分数(例如,过滤分数低于0.22的记录)并将计算的标签数组用作提示词。对于图像编辑或ControlNet类模型,可以利用OCR标签来筛选包含文本的图像进行训练,或使用人物相关标签来构建人像或姿态条件化的模型适配器。
ramanv-image-foundation is a production-grade, highly scalable dataset engineering platform specifically designed for processing and constructing large-scale datasets required for training and fine-tuning text-to-image generation models (e.g., FLUX, SDXL, and image editing models). The platform is designed to handle over 10 million images, and supports a range of functions including multi-GPU image caption generation, optical character recognition (OCR) text extraction, CLIP image-text alignment scoring, LAION aesthetic scoring, duplicate detection, quality filtering, and automatic upload to the Hugging Face Hub. The core content of the dataset is organized into 86 clearly defined category folders, covering a wide range of thematic domains, including Design and Layout (e.g., marketing materials, brand identities, posters, advertisements, icons, illustrations, 3D content, packaging, social media assets), Business and Products (e.g., electronics, fashion, cosmetics, food), Vehicles (e.g., cars, motorcycles, airplanes, vessels), Space and Architecture (e.g., architecture, real estate, interior design, office spaces, travel, hospitality), Nature and Landscapes (e.g., landscapes, forests, mountains, beaches, rivers, outer space), Science and Technology, Culture and People (e.g., festivals, temples, weddings, urban street scenes, portrait photography), and Documents and Text (e.g., receipts, tables, books, menus, newspapers, signage, certificates, handwritten content). The data is stored and provided in multiple formats: raw image files (.jpg/.png) organized by category; raw descriptive text files; metadata JSON files conforming to a specific schema (containing 49 fields); dataset manifest files for training, validation, and testing (in JSON Lines format); pre-computed CLIP/SigLIP feature vectors; and generated quality assessment reports and preview thumbnails. The data processing workflow is standardized, including caption generation, OCR extraction, alignment and aesthetic score calculation, duplicate removal and quality filtering, as well as dataset partitioning. This dataset is particularly suitable for training state-of-the-art text-to-image models that require high-quality, diverse visual content and rich textual descriptions. For the FLUX model, it is recommended to utilize detailed long captions, perform training with aspect ratio bucketing, and apply aesthetic score filtering (e.g., filtering out images with scores below 6.0). For the SDXL model, it is recommended to use CLIP alignment scores (e.g., filtering out records with scores below 0.22) and use the computed label arrays as prompts. For image editing or ControlNet-style models, OCR labels can be used to filter images containing text for training, or person-related labels can be used to build portrait or pose-conditioned model adapters.
数据集概述:ramanv-image-foundation
基本信息
- 名称:ramanv-image-foundation
- 许可证:Apache License 2.0
- 地址:https://huggingface.co/datasets/lingamvamshikrishnareddy/ramanv-image-foundation
数据集规模与能力
- 可处理高达 1000万+ 张图像
- 支持多GPU标注、OCR提取、CLIP对齐评分、美学分析、去重检查、质量过滤
- 支持自动上传至Hugging Face Hub
数据组织与结构
文件夹布局
数据集包含以下主要目录:
- images/:按86个类别分组的图像文件夹
- captions/:原始描述文本文件
- metadata/:包含49个字段、符合Schema的JSON文件
- manifests/:数据集清单文件(train.jsonl、validation.jsonl、test.jsonl、all.jsonl)
- quality/:质量指标与报告
- splits/:划分指标与记录
- embeddings/:预计算的CLIP/SigLIP特征矩阵(.npy格式)
- thumbnails/:生成的预览缩略图
- configs/:流水线配置文件(YAML格式)
- docs/:详细指南文档
- logs/:轮转日志输出文件
- src/:核心工具代码库
图像类别(86个类别)
数据集将资源组织为86个不同类别文件夹,包括:
- 设计与布局:市场营销、品牌、海报、传单、横幅、名片、广告、标志、图标、插图、3D、包装、模型、社交媒体
- 商业:产品、电子产品、时尚、化妆品、食物、餐厅
- 交通:车辆、汽车、摩托车、公交车、火车、飞机、船只
- 空间:建筑、房地产、室内、住宅、办公室、工作区、旅行、酒店
- 自然与科学:风景、森林、山脉、海滩、河流、太空、科学、技术、医疗、教育、体育
- 文化与人物:节日、寺庙、婚礼、印度、村庄、城市、街道、人类、肖像
- 文档与文字:文档、收据、表格、书籍、菜单、报纸、排版、招牌、证书、白板、手写
数据流水线处理流程
- 使用VLM生成标注
- 提取OCR文本
- 计算CLIP对齐度并提取嵌入特征
- 计算LAION美学评分
- 过滤低质量/模糊图像并去重
- 划分训练/验证/测试集
- 验证数据Schema
模型训练指南
微调FLUX模型
- 长标注:利用
long_caption元数据字段提供完整场景描述 - 宽高比分桶:按计算的
aspect_ratio分组训练条目 - 美学过滤:排除
aesthetic_score < 6.0的图像
训练SDXL模型
- CLIP分数:过滤
clip_score < 0.22的记录 - 标签:使用计算的
tags数组作为条件提示
训练编辑/ControlNet模型
- 使用OCR标签(
contains_text: true)隔离文本密集型图形 - 通过
contains_people或people_count进行筛选,构建肖像或姿态条件适配器





