遇见数据集

VoicePA

收藏
Mendeley Data2024-05-10 更新2024-06-28 收录
数据链接:
官方服务:

资源简介:

VoicePA is a dataset for speaker recognition and voice presentation attack detection (anti-spoofing).The dataset contains a set of Bona Fide (genuine) voice samples from 44 speakers and 24 different types of speech presentation attacks (spoofing attacks). The attacks were created using the Bona Fide data recorded for the AVSpoof dataset. Genuine data The genuine (non-attack) data is taken directly from 'AVspoof' database and can be used by both automatic speaker verification (ASV) and presentation attack detection (PAD) systems (folder 'genuine' contains this data). The genuine data acquisition process lasted approximately two months with 44 subjects, each participating in four different sessions configured in different environmental setups. During each recording session, subjects were asked to speak out prepared (read) speech, pass-phrases and free speech recorded with three devices: one laptop with high-quality microphone and two mobile phones (iPhone 3GS and Samsung S3). Attack data Based on the genuine data, 24 types of presentation attacks were generated. Attacks were recorded in 3 different environments (two typical offices and a large conference room), using 5 different playback devices, including built-in laptop speakers, high quality speakers, and three phones: iPhone 3GS, iPhone 6S, and Samsung S3, and assuming an ASV system running on either laptop, iPhone 3GS, or Samsung S3. In addition to a replay type of attacks (speech is recorded and replayed to the microphone of an ASV system), two types of synthetic speech were also replayed: speech synthesis and voice conversion (for the details on these algorithms, please refer to the paper below published in BTAS 2015 and describing 'AVspoof' database). Protocols The data in 'voicePA' database is split into three non-overlapping subsets: training (genuine and attack samples from 4 female and 10 male subjects), development or 'Dev' (genuine and attack samples from 4 female and 10 male subjects), and evaluation or 'Eval' (genuine and attack samples from 5 female and 11 male subjects). Reference Pavel Korshunov, André R. Goncalves, Ricardo P. V. Violato, Flávio O. Simões and Sébastien Marcel. "On the Use of Convolutional Neural Networks for Speech Presentation Attack Detection", International Conference on Identity, Security and Behavior Analysis, 2018. 10.1109/ISBA.2018.8311474 http://publications.idiap.ch/index.php/publications/show/3779

VoicePA是一款面向说话人识别与语音呈现攻击检测(反欺骗)的数据集。该数据集包含来自44位说话人的真实(Bona Fide)语音样本,以及24种不同类型的语音呈现攻击(欺骗攻击)样本。此类攻击样本均基于为AVSpoof数据集录制的真实语音数据生成。 真实数据:真实(非攻击)数据直接取自AVSpoof数据库,可同时用于自动说话人验证(Automatic Speaker Verification, ASV)与呈现攻击检测(Presentation Attack Detection, PAD)系统,此类数据存储于'genuine'文件夹中。真实数据的采集过程历时约两个月,共有44名受试者参与,每名受试者需在四种不同环境配置下完成四次录制。每次录制过程中,受试者需朗读预制文本、口令短语以及自由语音,录制设备包含一台搭载高品质麦克风的笔记本电脑与两部移动设备(iPhone 3GS及三星S3)。 攻击数据:基于上述真实语音数据,研究人员生成了24种语音呈现攻击样本。攻击样本在三种不同环境(两间典型办公室与一间大型会议室)中录制,使用了五种不同的播放设备:笔记本电脑内置扬声器、高品质扬声器,以及三款移动设备(iPhone 3GS、iPhone 6S与三星S3),且模拟了运行于笔记本电脑、iPhone 3GS或三星S3之上的ASV系统场景。除重播类攻击(将语音录制后重播至ASV系统的麦克风)外,本次数据集还包含两类合成语音重播攻击:语音合成与语音转换(相关算法细节可参阅2015年BTAS会议上发表的、描述AVSpoof数据库的相关论文)。 数据划分协议:VoicePA数据库的数据被划分为三个互不重叠的子集:训练集(取自4名女性与10名男性受试者的真实与攻击样本)、开发集(简称Dev集,取自4名女性与10名男性受试者的真实与攻击样本),以及评估集(简称Eval集,取自5名女性与11名男性受试者的真实与攻击样本)。 参考文献:Pavel Korshunov、André R. Goncalves、Ricardo P. V. Violato、Flávio O. Simões 与 Sébastien Marcel. 《基于卷积神经网络的语音呈现攻击检测应用》,国际身份、安全与行为分析会议(International Conference on Identity, Security and Behavior Analysis, ISBA 2018),2018. DOI: 10.1109/ISBA.2018.8311474 链接:http://publications.idiap.ch/index.php/publications/show/3779

创建时间:
2023-06-28
搜集汇总
数据集介绍
VoicePA 数据集图片
背景与挑战
背景概述
VoicePA数据集包含44位说话人的真实语音样本和24种语音呈现攻击样本,适用于说话人识别和反欺骗研究。数据集分为训练、开发和评估三个子集,支持多种设备和环境下的语音采集和攻击模拟。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务