IndieFake Dataset (IFD)
收藏资源简介:
IndieFake数据集(IFD)是一个包含印度英语说话者真实和深度伪造音频的基准数据集。它旨在解决现有数据集中缺乏南亚口音的问题,并包含50名英语说印度人的27.17小时的真实和深度伪造音频。数据集包含11.3小时的真实音频样本和15.82小时的深度伪造音频样本,平均音频样本长度为5秒。IFD具有平衡的数据分布,并包括说话人级别的特征描述,这在ASVspoof21(DF)等数据集中是缺失的。该数据集已被评估用于音频深度伪造检测,并与现有的ASVspoof21(DF)和In-The-Wild(ITW)数据集进行了比较,证明了其有效性。
IndieFake Dataset (IFD) is a benchmark dataset containing genuine and deepfake audio from Indian English speakers. It aims to address the lack of South Asian accents in existing datasets, and comprises 27.17 hours of genuine and deepfake audio from 50 Indian English speakers. The dataset includes 11.3 hours of genuine audio samples and 15.82 hours of deepfake audio samples, with an average audio sample length of 5 seconds. IFD features a balanced data distribution and includes speaker-level feature descriptions, which are absent in datasets such as ASVspoof21 (DF). This dataset has been evaluated for audio deepfake detection, and benchmarked against existing datasets including ASVspoof21 (DF) and In-The-Wild (ITW), which validates its effectiveness.




