So-Fake-Set
收藏资源简介:
So-Fake-Set是一个用于社交媒体图像伪造检测的大规模、多样化的数据集。它包含输入图像(包括真实图像、全合成图像和篡改图像)、二值掩模(突出显示篡改图像中的操作区域)、分类标签、生成器信息(对于真实图像此字段为None)、原始文件名和数据集分割(训练集和验证集)。
So-Fake-Set is a large-scale and diverse dataset dedicated to social media image forgery detection. It includes input images (comprising real images, fully synthetic images, and tampered images), binary masks that highlight the manipulated regions within tampered images, classification labels, generator information (this field is set to None for real images), original filenames, and dataset splits (training set and validation set).
So-Fake-Set 数据集概述
数据集基本信息
- 数据集名称: So-Fake-Set
- 项目主页: https://hzlsaber.github.io/projects/So-Fake/
- 代码仓库: https://github.com/hzlsaber/So-Fake
- 联系人: Zhenglin Huang (zhenglin@liverpool.ac.uk)
数据集简介
So-Fake-Set是一个大规模、多样化的社交媒体图像伪造检测数据集。
数据集结构
数据特征
- image (图像类型): 输入图像,包括真实图像、全合成图像和篡改图像
- mask (图像类型): 二进制掩码,突出显示篡改图像中的 manipulated 区域
- label (字符串类型): 分类类别
- generator (字符串类型): 生成器/源模型,真实图像此字段为
None - filename (字符串类型): 图像的原始文件名
- split (字符串类型): 训练集和验证集
数据划分
- 训练集: 1,990,070个样本,1,171,356,577,479字节
- 验证集: 235,670个样本,115,612,108,632字节
存储信息
- 下载大小: 1,281,505,074,183字节
- 数据集大小: 1,286,968,686,111字节
许可信息
本作品采用知识共享署名4.0国际许可协议。
引用信息
如需使用本数据集,请引用以下论文:
@misc{huang2025sofakebenchmarkingexplainingsocial, title={So-Fake: Benchmarking and Explaining Social Media Image Forgery Detection}, author={Zhenglin Huang and Tianxiao Li and Xiangtai Li and Haiquan Wen and Yiwei He and Jiangning Zhang and Hao Fei and Xi Yang and Xiaowei Huang and Bei Peng and Guangliang Cheng}, year={2025}, eprint={2505.18660}, archivePrefix={arXiv}, url={https://arxiv.org/abs/2505.18660}, }




