five

SWAN-DF database of audio-video deepfakes|深度伪造数据集|音频-视频处理数据集

收藏
Mendeley Data2024-05-10 更新2024-06-28 收录
深度伪造
音频-视频处理
下载链接:
https://zenodo.org/records/8365616
下载链接
链接失效反馈
资源简介:
Description SWAN-DF: the first high fidelity publicly available dataset of realistic audio-visual deepfakes, where both faces and voices appear and sound like the target person. The SWAN-DF dataset is based on the public SWAN database of real videos recorded in HD on iPhone and iPad Pro (in year 2019). For 30 pairs of manually selected people from SWAN, we swapped faces and voices using several autoencoder-based face swapping models and using several blending techniques from the well-known open source repo DeepFaceLab and voice conversion (or voice cloning) methods, including zero-shot YourTTS, DiffVC, HiFiVC, and several models from FreeVC. For each model and each blending technique, there are 960 video deepfakes. We used three types of models of the following resolutions: 160x160, 256x256, and 320x320 pixels. We took one pre-trained model corresponding for each resolution, and tuned it for each of the 30 pairs (both ways) of subjects for 50K iterations. Then, when generating deepfake videos for each pair of subjects, we used one of the tuned models and a way to blend the generated image back into the original frame, which we call blending technique. SWAN-DF dataset contains 25 different combinations of models and blending, which means the total number of deepfake videos is 960*25=24000. We generated speech deepfakes using four voice conversion methods: YourTTS, HiFiVC, DiffVC, and FreeVC. We did not use text to speech methods for our video deepfakes, since the speech they produce is not synchronized with the lip movements in the video. For YourTTS, HiFiVC, and DiffVC methods, we used the pretrained models provided by the authors. HiFiVC was pretrained on VCTK, DiffVC on LibriTTS, and YourTTS on both VCTK and LibriTTS datasets. For FreeVC, we generated audio deepfakes for several variants: using the provided pretrained models (for 16Hz with and without pretrained speaker encoder and for 24Hz with pretrained speaker encoder) as is and by tuning 16Hz model either from scratch or starting from the pretrained version for different number of iterations on the mixture of VCTK and SWAN data. In total, SWAN-DF contains 12 different variations of audio deepfakes: one for each of YourTTS, HiFiVC, and DiffVC and 9 variants of FreeVC. Acknowledgements If you use this database, please cite the following publication: Pavel Korshunov, Haolin Chen, Philip N. Garner, and Sébastien Marcel, "Vulnerability of Automatic Identity Recognition to Audio-Visual Deepfakes", IEEE International Joint Conference on Biometrics (IJCB), September 2023. https://publications.idiap.ch/publications/show/5092
创建时间:
2023-09-25
用户留言
有没有相关的论文或文献参考?
这个数据集是基于什么背景创建的?
数据集的作者是谁?
能帮我联系到这个数据集的作者吗?
这个数据集如何下载?
点击留言
数据主题
具身智能
数据集  4099个
机构  8个
大模型
数据集  439个
机构  10个
无人机
数据集  37个
机构  6个
指令微调
数据集  36个
机构  6个
蛋白质结构
数据集  50个
机构  8个
空间智能
数据集  21个
机构  5个
5,000+
优质数据集
54 个
任务类型
进入经典数据集
热门数据集

CHCrack5K

CHCrack5K是一个用于高级裂缝检测研究的强大数据集。它将11个公开的裂缝数据集整合为一个统一的数据集,包含5,014个标记图像样本。每个数据集都经过特定的预处理,以将所有样本标准化为480×480像素的分辨率。该数据集提供了多种裂缝结构,为测试稳健的裂缝检测算法提供了更具挑战性和现实性的基准。

github 收录

中国1km分辨率逐月降水量数据集(1901-2024)

该数据集为中国逐月降水量数据,空间分辨率为0.0083333°(约1km),时间为1901.1-2024.12。数据格式为NETCDF,即.nc格式。该数据集是根据CRU发布的全球0.5°气候数据集以及WorldClim发布的全球高分辨率气候数据集,通过Delta空间降尺度方案在中国降尺度生成的。并且,使用496个独立气象观测点数据进行验证,验证结果可信。本数据集包含的地理空间范围是全国主要陆地(包含港澳台地区),不含南海岛礁等区域。为了便于存储,数据均为int16型存于nc文件中,降水单位为0.1mm。 nc数据可使用ArcMAP软件打开制图; 并可用Matlab软件进行提取处理,Matlab发布了读入与存储nc文件的函数,读取函数为ncread,切换到nc文件存储文件夹,语句表达为:ncread (‘XXX.nc’,‘var’, [i j t],[leni lenj lent]),其中XXX.nc为文件名,为字符串需要’’;var是从XXX.nc中读取的变量名,为字符串需要’’;i、j、t分别为读取数据的起始行、列、时间,leni、lenj、lent i分别为在行、列、时间维度上读取的长度。这样,研究区内任何地区、任何时间段均可用此函数读取。Matlab的help里面有很多关于nc数据的命令,可查看。数据坐标系统建议使用WGS84。

国家青藏高原科学数据中心 收录

中国交通事故深度调查(CIDAS)数据集

交通事故深度调查数据通过采用科学系统方法现场调查中国道路上实际发生交通事故相关的道路环境、道路交通行为、车辆损坏、人员损伤信息,以探究碰撞事故中车损和人伤机理。目前已积累深度调查事故10000余例,单个案例信息包含人、车 、路和环境多维信息组成的3000多个字段。该数据集可作为深入分析中国道路交通事故工况特征,探索事故预防和损伤防护措施的关键数据源,为制定汽车安全法规和标准、完善汽车测评试验规程、

北方大数据交易中心 收录

LUNA16

LUNA16(肺结节分析)数据集是用于肺分割的数据集。它由 1,186 个肺结节组成,在 888 次 CT 扫描中进行了注释。

OpenDataLab 收录

UIEB, U45, LSUI

本仓库提供了水下图像增强方法和数据集的实现,包括UIEB、U45和LSUI等数据集,用于支持水下图像增强的研究和开发。

github 收录