MVAD
收藏资源简介:
MVAD是首个专门为检测AI生成的多模态视频-音频内容而设计的通用数据集。它涵盖两个视觉领域(写实和动漫风格)和四个主要类别(人类、动物、物体和场景),包含三种视频-音频伪造类型和四种模态组合(假-假、假-真、真-假、真-真)。数据集包含205,758个多模态视频-音频样本,使用超过20种不同方法生成,其中104,578个伪造样本和101,000个真实样本,伪造与真实样本比例为1:1。
MVAD is the first general-purpose dataset specifically designed for detecting AI-generated multimodal video-audio content. It covers two visual domains (realistic and anime styles) and four main categories (humans, animals, objects, and scenes), and includes three types of video-audio forgeries as well as four modality combinations (fake-fake, fake-real, real-fake, real-real). The dataset contains 205,758 multimodal video-audio samples generated by over 20 different methods, among which there are 104,578 forged samples and 101,000 real samples, with a 1:1 ratio between forged and real samples.




