A2IR: Audio-to-Image Representation
收藏资源简介:
A2IR is a dataset for synthetic audio detection using deep learning. It includes five audio-to-image representations for natural and synthetic audio: spectrograms, histograms, scatter plots, bispectrum phase plots and bispectrum magnitude plots. Each category is divided into 3 subsets: training 56.72% (11,400 images), validation 33.83% (6,800 images), and test 9.47% (1,900 images). In each subset, the images are separated into two folders, natural and synthetic, with a balanced classification (i.e. each class has the same number of images as the other or very similar).
A2IR是一款用于深度学习合成音频检测的数据集。该数据集包含五种面向自然音频与合成音频的音频转图像表征形式:频谱图(spectrograms)、直方图(histograms)、散点图(scatter plots)、双谱相位图(bispectrum phase plots)以及双谱幅度图(bispectrum magnitude plots)。整个数据集的图像被划分为三个子集:训练集占比56.72%(共计11400张图像)、验证集占比33.83%(共计6800张图像)以及测试集占比9.47%(共计1900张图像)。在每个子集内部,图像被分为自然音频和合成音频两个文件夹,且分类平衡(即两类图像的数量完全相等或极为接近)。



