遇见数据集

DG-X SMFD v1.0: DeepGuard-X Synthetic Multimodal Forensic Dataset

收藏
Zenodo2026-08-04 更新2026-08-13 收录
官方服务:

资源简介:

A literature-calibrated synthetic dataset for developing and benchmarking multi-modal (image + video + audio) manipulated-media detection pipelines. IMPORTANT: This is a fully synthetic dataset. It contains no real images, video, or audio, and does not depict, represent, or derive from any real individual. Do not cite or use this dataset as if it contains genuine forensic or biometric media. Contents: Two CSV files, 50,000 cases each —- SYNTHETIC_dgx_smfd_independent_v1.csv — baseline condition, per-modality artifact features generated independently- SYNTHETIC_dgx_smfd_correlated_v1.csv — stress-test condition, artifact features share a per-case latent "manipulation quality" factor (correlation strength 0.6) Each file includes a fixed train/validation/test split (70/15/15) and 15 per-modality artifact features (5 image, 5 video, 5 audio), whose real-vs-fake separability is set analytically from published detection-accuracy statistics (AUC) for each artifact family, via d = √2·Φ⁻¹(AUC). Realistic evidence gaps are simulated: ~3.0% of cases lack image evidence, ~21.9% lack video, ~35.2% lack audio. Intended use: pipeline development, methodology benchmarking, and reproducible research on fusion, uncertainty-quantification, and explainability architectures for manipulated-media detection — not as a substitute for evaluation on genuine deepfake corpora before real-world deployment claims. Full schema, generation methodology, and licensing details are included in the accompanying README.md and LICENSE.txt. Accompanying paper: "DeepGuard-X: An Explainable, Uncertainty-Aware Multi-Modal Fusion Architecture for Deepfake and Manipulated Media Triage" (2026).

提供机构:
Zenodo
创建时间:
2026-08-04
二维码
社区交流群
二维码
科研交流群
商业服务