遇见数据集

The Digital Forensics 2023 dataset - DF2023

收藏
Zenodo2025-04-03 更新2026-05-25 收录
官方服务:

资源简介:

For a detailed description of the DF2023 dataset, please refer to: @inproceedings{Fischinger2023DFNet, title={DF2023: The Digital Forensics 2023 Dataset for Image Forgery Detection}, author={David Fischinger and Martin Boyer}, journal={The 25th Irish Machine Vision and Image Processing conference. (IMVIP)}, year={2023} } DF2023 is a dataset for image forgery detection and localization. The training and validation datasets contain 1,000,000/5,000 manipulated images (and the ground truth masks). The DF2023 training dataset comprises: 100K forged images produced by removal (inpainting) operations 200K images produced by enhancement modifications 300K copy-move manipulated images and 400K spliced images === Naming convention === The naming convention of DF2023 encodes information about the applied manipulations. Each image name has the following form: COCO_DF_0123456789_NNNNNNNN.{EXT} (e.g. COCO_DF_E000G40117_00200620.jpg) After the identifier of the image data source ("COCO") and the self-reference to the Digital Forensics ("DF") dataset, there are 10 digits as placeholders for the manipulation. Position 0 defines the manipulation types copy-move, splicing, removal, enhancement ([C,S,R,E]). The following digits 1-9 represent donor patch manipulations. For positions [1,2,7,8] (resample, flip, noise and brightness), a binary value indicates if this manipulation was applied to the donor image patch. Position 3 (rotate) indicates by the values 0-3 if the rotation was executed by 0, 90, 180 or 270 degrees. Position 4 defines if BoxBlur (B) or GaussianBlur (G) was used. Position 5 specifies the blurring radius. A value of 0 indicates that no blurring was executed. Position 6 indicates which of the Python-PIL contrast filters EDGE ENHANCE, EDGE ENHANCE MORE, SHARPEN, UnsharpMask or ImageEnhance (values 1-5) was applied. If none of them was applied, this value is set to 0. Finally, position 9 is set to the JPEG compression factor modulo 10, a value of 0 indicates that no JPEG compression was applied. The 8 characters NNNNNNNN in the image name template stand for a running number of the images. === Terms of Use / Licence === The DF2023 dataset is based on the MS COCO dataset. Therefore, rules for using the images form MS COCO apply also for DF2023: Images The COCO Consortium does not own the copyright of the images. Use of the images must abide by the Flickr Terms of Use. The users of the images accept full responsibility for the use of the dataset, including but not limited to the use of any copies of copyrighted images that they may create from the dataset.

如需获取DF2023(Digital Forensics 2023 Dataset,数字取证2023数据集)的详细说明,请参阅以下文献: @inproceedings{Fischinger2023DFNet, title={DF2023: 面向图像篡改检测的数字取证2023数据集}, author={David Fischinger 与 Martin Boyer}, conference={第25届爱尔兰机器视觉与图像处理大会(IMVIP)}, year={2023} } DF2023是一款面向图像篡改检测与定位的数据集。其训练集与验证集分别包含100万张与5000张经过篡改的图像(附带真值掩码)。 DF2023训练集包含如下四类数据: 1. 10万张通过移除(图像修复,inpainting)操作生成的伪造图像 2. 20万张经过增强修改的图像 3. 30万张复制-粘贴篡改图像 4. 40万张拼接篡改图像 === 命名规则 === DF2023的命名规则对所应用的篡改操作进行了编码。每张图像的命名格式如下: COCO_DF_0123456789_NNNNNNNN.{EXT}(例如:COCO_DF_E000G40117_00200620.jpg) 在图像数据源标识符"COCO"与数字取证数据集自标识"DF"之后,有10位数字作为篡改操作的占位符。第0位用于定义篡改类型:复制-移动(C)、拼接(S)、移除(R)、增强(E)。后续第1至9位则代表供体图像块的篡改操作: - 第[1,2,7,8]位分别对应重采样、翻转、噪声与亮度调整,采用二进制值表示是否对供体图像块应用了该操作; - 第3位(旋转)通过0-3的数值分别代表旋转角度为0°、90°、180°或270°; - 第4位用于指定使用的模糊类型:方框模糊(BoxBlur)或高斯模糊(GaussianBlur); - 第5位指定模糊半径,取值为0则表示未应用模糊操作; - 第6位表示应用的Python-PIL对比度滤镜类型:边缘增强(EDGE ENHANCE)、强边缘增强(EDGE ENHANCE MORE)、锐化(SHARPEN)、非锐化掩模(UnsharpMask)或图像增强(ImageEnhance),对应取值1-5;若未应用此类滤镜,则该位取值为0; - 第9位为JPEG(Joint Photographic Experts Group)压缩系数模10的结果,取值为0则表示未应用JPEG压缩。 命名格式中的8位字符NNNNNNNN代表图像的流水号。 === 使用条款 / 许可协议 === DF2023数据集基于MS COCO(Microsoft COCO,微软COCO)数据集构建,因此MS COCO数据集的图像使用规则同样适用于DF2023: #### 图像版权说明 COCO联盟并不拥有这些图像的版权。使用这些图像必须遵守Flickr服务条款。数据集使用者需对数据集的使用承担全部责任,包括但不限于因使用本数据集生成的任何受版权保护的图像副本所带来的相关责任。

提供机构:
Zenodo
创建时间:
2022-11-03
二维码
社区交流群
二维码
科研交流群
商业服务