Irish-Accented English Audio-Visual Deepfake Datasets with Deep Packet Inspection-Inspired Media Integrity Validation
收藏资源简介:
This dataset contains Irish-accented English audio-only and synchronised audio–video samples curated for research on deepfake detection, multimodal learning, media integrity validation, cybersecurity, accent robustness, and bias-aware evaluation. This repository contains both the main classification dataset and a gender-based subset organised for speaker-gender-aware analysis. The deposit includes authentic and synthetic samples, file-level metadata, labels, source-provenance documentation, and validation materials. Authentic media were retrieved from publicly accessible Archive.org item pages and manually reviewed for Irish-accented English speech. Synthetic samples were generated using generic text-to-speech and prompt-based media generation workflows. No synthetic sample was generated to clone, impersonate, face-swap, lip-sync, or reproduce the voice, face, likeness, or identity of any known individual. The authors do not claim ownership over third-party authentic Archive.org media and do not relicense those third-party media. The authors’ licence applies to the metadata, labels, documentation, validation scripts, processing code, and author-generated synthetic media. Third-party authentic media remain subject to their original rights and applicable Archive.org item-level terms.



