遇见数据集

ManipFace-XAI: A Synthetic Manipulated Face Dataset for Explainable Deepfake Analysis

收藏
Zenodo2026-05-30 更新2026-06-05 收录
官方服务:

资源简介:

This dataset contains the generated fake face image subset created to support the development and evaluation of hierarchical deepfake detection models, where the first stage performs real/fake classification and the second stage performs fake-type attribution. The released subset includes only fake face images generated using three manipulation or synthesis approaches: diffusion-based generation, SimSwap face swapping, and StyleGAN2 synthetic face generation. The original experimental dataset also included real face images, but these images are not redistributed in this Zenodo record. In addition, direct references to the original real image paths or source images were removed from the released metadata. The dataset combines fake images with different provenance. The diffusion and SimSwap subsets were created using images derived from an initial real/fake face dataset used in the experimental pipeline. This original dataset was based on the Deepfake and Real Images Dataset available on Kaggle and is associated with the OpenForensics dataset introduced by Le et al. (2021). By contrast, the StyleGAN2 subset contains fully synthetic faces generated independently from random seeds using a pretrained StyleGAN2 model, and therefore was not derived from the original real-image dataset. The final released fake subset contains 72,335 images, distributed as follows: Diffusion: 26,023 images SimSwap: 20,289 images StyleGAN2: 26,023 images The dataset is accompanied by a sanitized metadata file, split.csv, containing only the following fields: name, label, split, fake_type. The name column stores only the image filename, without local or absolute paths. The label column indicates that all released samples are fake. The split column indicates the experimental partition used in the project, and the fake_type column specifies the manipulation or generation category. The released split distribution is: Train: 50,623 images Validation: 10,849 images Test: 10,863 images This dataset is intended for academic research on deepfake detection, fake-type attribution, and explainable artificial intelligence methods applied to facial image analysis. It should not be used for impersonation, identity misuse, or any harmful synthetic media generation purposes. Original dataset reference: Le, T. N., Nguyen, H. H., Yamagishi, J., & Echizen, I. (2021). OpenForensics: Large-Scale Challenging Dataset for Multi-Face Forgery Detection and Segmentation In-The-Wild. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 10117–10127.

提供机构:
Zenodo
创建时间:
2026-05-30
二维码
社区交流群
二维码
科研交流群
商业服务