遇见数据集

FakeReveal: Multi-modal_Audio-Visual_feature-based dataset

收藏
Mendeley Data2026-09-08 收录
官方服务:

资源简介:

This dataset is a multimodal, feature-based dataset developed for an audio-visual deepfake detection system. It provides numerical feature representations for 7,602 samples, including 3,396 real and 4,206 fake samples. The AVLips dataset was used as the source of the video samples from which this feature-based dataset was derived. Its video samples were computationally processed through a multimodal pipeline combining visual speech and audio speech recognition. The audio was processed using Whisper (base) to generate an audio speech transcription. In parallel, the mouth region was extracted using MediaPipe Face Mesh and processed using the pre-trained AV-HuBERT base_vox_433h model to generate a visual speech transcription based on lip movements. These two textual representations were then compared to capture inconsistencies between the spoken audio and the lip movements. Three cross-modal textual comparison features were extracted from the Whisper and AV-HuBERT outputs. Before comparison, both transcripts were normalised by expanding contractions, converting text to lowercase, removing punctuation, and standardising whitespace. The obtained features include Word Error Rate (WER), text similarity score, and correct word count. Together, these features provide complementary measurements of word-level error, overall textual similarity, and exact lexical agreement between audio and visual speech. In a parallel audio-based pipeline, MFCC and spectral features were extracted using the Librosa library. The resulting 50-dimensional acoustic representation consists of 40 Mel-Frequency Cepstral Coefficients (MFCCs), 1 Spectral Centroid, 1 Spectral Roll-off, 1 Spectral Bandwidth, and 7 Spectral Contrast features. A linear Support Vector Machine (SVM) classifier was then applied to the extracted MFCC and spectral features. Three model-derived outputs were retained for each sample: the prediction label, the prediction probability and the decision score. The dataset contains two complementary files: FakeReveal_Master_Features.csv, which provides the consolidated feature table and FakeReveal_MFCC_Spectral_Features.pkl, which preserves the 50-dimensional MFCC and spectral feature vectors. The FakeReveal_Master_Features.csv contains the following columns: Video_ID: Identifier. label: class: label, where 0 represents real and 1 represents fake. similarity: textual similarity between the audio transcription and the visual speech transcription. WER: measuring the word-level discrepancy between the audio and visual speech transcriptions. correct_words_#: Number of words correctly matched between the audio and visual speech transcriptions. mfcc_prediction: Binary prediction produced by the MFCC-based SVM classifier mfcc_probability: Prediction probability produced by the MFCC-based SVM classifier. mfcc_svm_decision_score: decision function score representing the sample's position relative to the classifier's decision boundary.

创建时间:
2026-08-24
二维码
社区交流群
二维码
科研交流群
商业服务