SDFVD: Small-scale Deepfake Forgery Video Dataset
收藏资源简介:
Small-scale Deepfake Forgery Video Dataset (SDFVD) is a custom dataset consisting of real and deepfake videos with diverse contexts designed to study and benchmark deepfake detection algorithms. The dataset comprising of a total of 106 videos, with 53 original and 53 deepfake videos. Equal number of real and deepfake videos, ensures balance for machine learning model training and evaluation. The original videos were collected from Pexels: a well- known provider of stock photography and stock footage(video). These videos include a variety of backgrounds, and the subjects represent different genders and ages, reflecting a diverse range of scenarios. The input videos have been pre-processed by cropping them to a length of approximately 4 to 5 seconds and resizing them to 720p resolution, ensuring a consistent and uniform format across the dataset. Deepfake videos were generated using Remaker AI employing face-swapping techniques. Remaker AI is an AI-powered platform that can generate images, swap faces in photos and videos, and edit content. The source face photos for these swaps were taken from Freepik: is an image bank website provides contents such as photographs, illustrations and vector images. SDFVD was created due to the lack of availability of any such comparable small-scale deepfake video datasets. Key benefits of such datasets are: • In educational settings or smaller research labs, smaller datasets can be particularly useful as they require fewer resources, allowing students and researchers to conduct experiments with limited budgets and computational resources. • Researchers can use small-scale datasets to quickly prototype new ideas, test concepts, and refine algorithms before scaling up to larger datasets. Overall, SDFVD offers a compact but diverse collection of real and deepfake videos, suitable for a variety of applications, including research, security, and education. It serves as a valuable resource for exploring the rapidly evolving field of deepfake technology and its impact on society.
小规模深度伪造视频数据集(Small-scale Deepfake Forgery Video Dataset,简称SDFVD)是一款专属定制数据集,涵盖真实视频与深度伪造(Deepfake)视频,场景多元,旨在用于深度伪造检测算法的研究与基准评测。该数据集总计包含106段视频,其中原始真实视频与深度伪造视频各53段。真实与伪造视频数量均等,可为机器学习模型的训练与评估提供均衡的数据分布。原始真实视频采集自Pexels——全球知名的正版商用图库与视频素材提供商。这些视频涵盖多样背景,拍摄对象包含不同性别与年龄段,场景分布广泛且多元。输入视频均经过标准化预处理:裁剪至约4至5秒的时长,并调整分辨率至720p,确保数据集内所有视频格式统一规范。深度伪造视频通过Remaker AI平台的人脸交换技术生成。Remaker AI是一款人工智能驱动的内容创作平台,支持图像生成、照片与视频人脸交换以及内容编辑等功能。本次人脸交换所用的源面部照片采集自Freepik——一家提供摄影作品、插画与矢量图像等素材的在线图库网站。鉴于当前尚无同类可比的小规模深度伪造视频数据集,本数据集SDFVD应运而生。此类小规模数据集的核心优势在于: • 对于教学场景或小型研究实验室而言,小型数据集所需资源更少,可帮助学生与研究人员在预算有限、计算资源不足的情况下开展实验。 • 研究人员可借助小规模数据集快速验证新想法、测试研究概念并优化算法,后续再将方案拓展至大规模数据集。总体而言,SDFVD提供了一个体量小巧但场景多元的真实与深度伪造视频集合,可适用于研究、安全与教育等多类应用场景,可为快速发展的深度伪造技术领域及其社会影响研究提供宝贵的资源支撑。




