SGData
收藏资源简介:
SGData是一个面向语义感知的虚假图像/视频检测数据集,与SGDFD方法一同提出。该数据集涵盖来自ILSVRC2012的1000个语义类别,每个类别的真实图像由七种最先进的文本到图像生成器(StableDiffusion 3.5、FLUX.2、PixelArt-Σ、Kandinsky 5.0、HiDream-I1、Lumina-Image 2.0、Qwen-Image)根据QWEN3-VL自动生成的文本提示一对一地合成伪造图像。此外,还引入了额外的人脸语义类别,其中真实图像来自FaceCaptionHQ-4M,真实视频来自CelebV-Text,伪造人脸视频由三种视频生成模型(SANA-Video、HunyuanVideo 1.5、SkyReels-V2)生成。SGData的测试集中的人脸数据未经过对齐或中心裁剪处理,以模拟真实部署场景。数据集规模:训练集包含700,700张真实图像和700,700张伪造图像;验证集包含70,070张真实图像和70,070张伪造图像;测试集包含630,630张真实图像、630,630张伪造图像、600段真实视频和600段伪造视频。总计1,401,400张真实图像、1,401,400张伪造图像、600段真实视频和600段伪造视频。语义类别总数为1001,生成模型总数为10。数据组织方式为:真实图像和视频需用户自行从原始来源(ImageNet2012、FaceCaptionHQ-4M、CelebV-Text)下载,伪造图像和视频以tar包形式提供,按语义类别和生成器名称组织。
SGData is a semantic-aware fake image/video detection dataset proposed alongside the SGDFD method. It covers 1000 semantic categories from ILSVRC2012, with real images for each category and fake images synthesized by seven state-of-the-art text-to-image generators (StableDiffusion 3.5, FLUX.2, PixelArt-Σ, Kandinsky 5.0, HiDream-I1, Lumina-Image 2.0, Qwen-Image) based on text prompts automatically generated by QWEN3-VL. Additionally, an extra face semantic category is introduced, where real images are from FaceCaptionHQ-4M, real videos from CelebV-Text, and fake face videos are generated by three video generation models (SANA-Video, HunyuanVideo 1.5, SkyReels-V2). Face data in the test set of SGData are not aligned or center-cropped to simulate real deployment scenarios. Dataset scale: training set contains 700,700 real images and 700,700 fake images; validation set contains 70,070 real images and 70,070 fake images; test set contains 630,630 real images, 630,630 fake images, 600 real videos, and 600 fake videos. Total: 1,401,400 real images, 1,401,400 fake images, 600 real videos, and 600 fake videos. There are 1001 semantic categories and 10 generative models. Data organization: real images and videos need to be downloaded by users from original sources (ImageNet2012, FaceCaptionHQ-4M, CelebV-Text), while fake images and videos are provided as tar packages organized by semantic category and generator name.
SGData 是一个用于伪造内容检测的语义感知数据集,支持图像与视频两类媒体,总计包含约 280 万张图像与 1200 个视频。真实图像覆盖 ILSVRC2012 的全部 1000 个语义类别,并额外引入“人脸”类别(共 1001 类);虚假图像由 7 种先进的文生图模型按类别一对一生成,虚假视频由 3 种视频合成模型生成,且测试集中的人脸图像刻意不做对齐或中心裁剪,以贴近真实应用场景。数据规模按子集划分:训练集含真实/虚假图像各 700,700 张;验证集各 70,070 张;测试集各 630,630 张,并含真实/虚假视频各 600 段。生成模型方面,图像来自 StableDiffusion 3.5、FLUX.2、PixelArt-Σ、Kandinsky 5.0、HiDream-I1、Lumina-Image 2.0、Qwen-Image(每个模型对 1000 个类别各生成 200 张,并对人脸类别额外生成 200 张);视频来自 SANA-Video、HunyuanVideo 1.5、SkyReels-V2(各生成 200 段人脸视频,条件来源于 CelebV-Text)。数据集真实媒体需用户自行获取:真实图像来源于 ImageNet2012 与 FaceCaptionHQ-4M,需同意相应访问条款;真实视频来源于 CelebV-Text,亦需同意其条款。仓库结构为:真实图像(需自行下载)、fake 目录下按语义类别存放生成器压缩包、真实视频(需自行下载)、fake_videos 目录下存放各生成器的视频压缩包。数据集许可为 MIT,总规模在 1 亿至 10 亿条之间。




