DRAGON
收藏资源简介:
DRAGON数据集是迄今为止规模最大、最多样化的扩散模型检测和归因任务数据集,包含由25个不同的扩散模型生成的260万个合成图像,以及从ImageNet获取的133万个真实图像。数据集已预分为训练集和测试集,并组织为五个子集,每个子集包含不同数量的图像,以满足各种研究场景的需求。数据集的创建旨在支持研究人员开发用于分析可能由扩散模型生成的图像的方法,并解决数字取证社区近年来出现的重大挑战。此外,数据集还包含一个专门的测试集,旨在作为评估新开发方法的基准。
The DRAGON dataset is the largest and most diverse dataset to date for diffusion model detection and attribution tasks, containing 2.6 million synthetic images generated by 25 distinct diffusion models and 1.3 million real images sourced from ImageNet. The dataset is pre-split into training and test sets, and organized into five subsets with varying numbers of images to meet the needs of diverse research scenarios. It was developed to support researchers in developing methods for analyzing images potentially generated by diffusion models, and to address major challenges that have emerged in the digital forensics community in recent years. Additionally, the dataset includes a dedicated test set intended as a benchmark for evaluating newly developed methods.
DRAGON 数据集概述
基本信息
- 名称: DRAGON (Dataset of Realistic imAges Generated by diffusiON models)
- 许可证: Creative Commons Attribution Share Alike 4.0 International (cc-by-sa-4.0)
- 任务类别: 图像分类 (image-classification)
- 数据规模: 1M < n < 10M
- 数据集大小: 250万训练图像 + 10万测试图像
- 生成模型数量: 25种扩散模型
数据集描述
- 目的: 支持开发多媒体取证工具,专注于合成图像检测和模型归属任务
- 特点:
- 包含多样化主题的合成图像
- 提供多种规模子集(从XS到XL)
- 包含专门设计的测试集作为标准化基准
数据集结构
- 标注信息: 每张图像标注了生成模型和输入提示
- 生成方式:
- 基于1,000个ImageNet类别生成提示
- 每个模型每个提示生成100张训练图像和4张测试图像
子集规模
| 子集名称 | 训练图像数量 | 测试图像数量 | 提示数量 |
|---|---|---|---|
| ExtraSmall (XS) | 250 | 1,000 | 10 |
| Small (S) | 2,500 | 10,000 | 100 |
| Regular (R) | 25,000 | 10,000 | 100 |
| Large (L) | 250,000 | 100,000 | 1,000 |
| ExtraLarge (XL) | 2,500,000 | 100,000 | 1,000 |
文件配置
- ExtraSmall:
- 训练: train/xs/dragon_train_xs.tar
- 测试: test/dragon_test_00.tar
- Small:
- 训练: train/dragon_train_000.tar
- 测试: test/dragon_test_0?.tar
- Regular:
- 训练: train/dragon_train_00?.tar
- 测试: test/dragon_test_0?.tar
- Large:
- 训练: train/dragon_train_0??.tar
- 测试: test/dragon_test_??.tar
- ExtraLarge:
- 训练: train/dragon_train_???.tar
- 测试: test/dragon_test_??.tar
引用信息
bibtex @misc{bertazzini2025dragon, title={DRAGON: A Large-Scale Dataset of Realistic Images Generated by Diffusion Models}, author={Giulia Bertazzini and Daniele Baracchi and Dasara Shullani and Isao Echizen and Alessandro Piva}, year={2025}, eprint={2505.11257}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2505.11257}, }
联系方式
- Giulia Bertazzini: giulia.bertazzini@unifi.it
- Daniele Baracchi: daniele.baracchi@unifi.it




