DeepFake Videos Dataset - 5,000 Files for Facial Recognition and Computer Vision
收藏资源简介:
Overview DeepFake Videos Dataset is a commercial synthetic video dataset produced by Unidata, containing 5,000 video files featuring real people with AI-generated faces overlaid. Designed specifically for deepfake detection, facial recognition, and computer vision research, it provides realistic and diverse training material for building and benchmarking deepfake detectors and detection algorithms. Each video was created by generating fake faces and overlaying them onto authentic source videos of real individuals turning their heads in different directions. Dataset Composition - 5,000 video files, each corresponding to a unique person - 5,000 unique subjects with individual metadata records - AI face generation sources: aisaver.io, faceswapvideo.ai, magichour.ai - Data type: real videos with AI-generated faces overlaid - Labeling: metadata per subject — age, gender, ethnicity Subject Demographics Both male and female participants are represented in the dataset: - Gender breakdown: Male and Female subjects - Ethnicity breakdown: Asian (30%), African (70%) - Age range: Min = 18, Max = 80, Mean = 45 Each video is labeled with structured metadata covering age, gender, and ethnicity, enabling targeted filtering and stratified model training across demographic groups. Technical Specifications - Formats: MP4, MOV - Resolutions: 1920×1080p, 1280×720p, 720×480p, 640×480p, 480×360p, 1920×920p - Video duration: Mean = 9 sec, Median = 9 sec, Min = 2 sec, Max = 34 sec - Frames per second: Mean = 26.6 FPS - Recording devices: iPhone 13 (30%), Google Pixel (70%) Videos cover diverse individuals, backgrounds, and scenarios, making the dataset suitable for a wide range of detection methods and recognition systems. Data Collection Methodology The dataset was built by generating fake faces using generative AI models and overlaying them onto real video clips. This hybrid approach — combining real-world source videos with AI-generated faces — produces synthetic data that closely mimics actual deepfake content encountered in the wild, ensuring better results during model training and evaluation. Use Cases - Cybersecurity & Digital Forensics. Provides critical training data for developing deepfake detectors and detection algorithms. The dataset consists of both real videos and fake videos generated using advanced deepfake technology, allowing analysts to train detection systems capable of identifying synthetic media and protecting against identity fraud and misinformation. - AI & Machine Learning Research. Widely used in deep learning and machine learning projects. Models trained on this dataset achieve better accuracy in spotting AI-generated videos and distinguishing between real and fake content. The dataset consists of thousands of video clips with manually labelled metadata, supporting robust model development. - Media & Journalism. News organizations use such datasets to enhance video detection tools that verify YouTube videos, interviews, and shared clips. By training recognition systems on datasets containing both source videos and generated faces, journalists can validate footage and identify manipulated content. Compliance & Security The DeepFake Videos Dataset complies with GDPR and applicable data protection regulations. All data is stored on AWS cloud infrastructure certified to ISO 27001 and ISO 27701 standards, ensuring internationally recognized information security and privacy management. Summary The DeepFake Videos Dataset is a large-scale synthetic media collection purpose-built for deepfake detection research, facial recognition, and computer vision tasks. With 5,000 video clips covering 5,000 unique individuals, demographic metadata across age, gender, and ethnicity, multiple resolutions up to 1080p, and an average of 26.6 FPS — it delivers the diversity and scale needed to train detection systems that achieve better results against real-world deepfake content. Whether used for building detection algorithms, training deep learning models, or validating recognition systems, this dataset provides a comprehensive and legally compliant resource for teams working on synthetic media detection.



