OpenMMSec
收藏资源简介:
OpenMMSec数据集聚合了19个公共取证数据集的数据,涵盖10个真实世界数据集,包含超过330K样本,覆盖所有4个子领域(Deepfake、AIGC、IMDL、Doc)和98种图像伪造类型。它具有全面覆盖、大规模多样性、平衡分布、丰富真实图像来源和定位支持等关键特征。
The OpenMMSec dataset aggregates data from 19 public forensic datasets, covering 10 real-world datasets with over 330K samples, and spanning all 4 sub-fields (Deepfake, AIGC, IMDL, Doc) as well as 98 types of image forgeries. It features key characteristics including comprehensive coverage, large-scale diversity, balanced distribution, abundant real-world image sources, and localization support.
数据集概述:OpenMMSec
OpenMMSec 是一个大规模、多领域的伪造图像检测数据集,旨在支持统一的伪造图像检测模型研究。该数据集与 ICML 2026 论文《Can We Build a Monolithic Model for Fake Image Detection? SICA: Semantic-Induced Constrained Adaptation for Unified-Yet-Discriminative Artifact Feature Space Reconstruction》一同发布。
核心特性
- 全面覆盖:涵盖图像取证领域的 4 个主要子领域:Deepfake(深度伪造)、AIGC(AI生成内容)、IMDL(图像篡改检测与定位)、Doc(文档伪造)。
- 大规模与多样性:包含 333,583 张图像,覆盖 15 种主要伪造类型 和 98 种细粒度伪造类型。
- 平衡分布:在不同伪造类型之间仔细调整数据量,确保公平比较。
- 丰富的真实图像来源:真实图像来自 超过 10 个真实世界数据集。
- 定位支持:保留来自原始数据集(IMDL 和 Doc)的像素级掩码,以支持未来的定位研究。
数据集统计
| 数据分区 | 来源数据集数 | 主要类型数 | 细粒度类型数 | 真实图像数 | 伪造图像数 | 总计 |
|---|---|---|---|---|---|---|
| Deepfake | 6 | 4 | 45 | 29,000 | 65,636 | 94,636 |
| AIGC | 3 | 6 | 26 | 46,000 | 46,048 | 91,048 |
| IMDL | 7 | 3 | 9 | 47,914 | 51,000 | 98,914 |
| Doc | 3 | 2 | 18 | 6,388 | 42,597 | 48,985 |
| 总计 | 19 | 15 | 98 | 129,302 | 204,281 | 333,583 |
- 泛化评估划分:将 98 种细粒度类型划分为 26 种用于训练(81,632 训练 / 8,240 验证),其余 72 种用于测试(243,711 测试)。
数据来源
OpenMMSec 汇聚了来自 19 个公开取证数据集 的数据,并跨越 超过 10 个真实世界数据集。
许可证与下载
- 许可证:Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)。
- 访问权限:仅限教育机构和非营利组织的研究人员。
- 下载链接:
- 百度网盘:https://pan.baidu.com/s/1NYB7obv_1G-ECRvOA4zBeA?pwd=vhxy
- Google Drive:https://drive.google.com/file/d/1_rQHxS8zlQlXZY_e_TbRPIPSQFMgCTOT/view?usp=sharing
相关资源
- 论文:https://arxiv.org/pdf/2602.06676
- 官方仓库:https://github.com/venus-guangjian/SICA_OpenMMSec
- 预训练权重:https://drive.google.com/drive/folders/109nJHqK-REXj5rvgpUOnzF4e0YPMBZbP?usp=sharing
- 训练框架:ForensicHub (https://github.com/scu-zjz/ForensicHub)
引用
如使用本数据集,请引用以下论文:
@article{du2026can, title={Can We Build a Monolithic Model for Fake Image Detection? SICA: Semantic-Induced Constrained Adaptation for Unified-Yet-Discriminative Artifact Feature Space Reconstruction}, author={Du, Bo and Ma, Xiaochen and Zhu, Xuekang and Yang, Zhe and Niu, Chaogun and Fang, Mingqi and Wang, Zhenming and Liu, Jingjing and Liu, Jian and Zhou, Ji-Zhe}, journal={arXiv preprint arXiv:2602.06676}, year={2026} }




