CompSpoof
收藏资源简介:
CompSpoof数据集是一个用于研究组件级音频反欺骗的公开数据集,由美国洛杉矶的OfSpectrum公司和中国昆山杜克大学数字创新研究中心的研究人员创建。该数据集包含约2500个音频样本,涵盖真实语音、欺骗语音、真实环境声和欺骗环境声的多种组合,用于训练和测试组件级音频欺骗检测模型。数据集分为训练集、开发集和评估集,用于支持音频欺骗检测研究,特别是针对组件级欺骗场景。
The CompSpoof dataset is a public dataset for research on component-level audio anti-spoofing, created by researchers from OfSpectrum Corporation based in Los Angeles, the United States, and the Digital Innovation Research Center of Duke Kunshan University in China. It contains approximately 2,500 audio samples, covering various combinations of genuine speech, spoofed speech, genuine ambient sounds and spoofed ambient sounds, which are used for training and testing component-level audio spoofing detection models. The dataset is divided into training set, development set and evaluation set to support audio spoofing detection research, especially for component-level spoofing scenarios.
CompSpoof Dataset 概述
数据集简介
CompSpoof数据集专为研究组件级反欺骗而设计,其中语音或环境声音组件(或两者)可能被欺骗。
数据集概览
- 总样本数: 2,500
- 类别数: 5(每类500个样本)
- 时长: 5–21秒
- 采样率: 16 kHz
- 数据划分: 70%训练集、10%开发集、20%评估集(分层划分以保持类别平衡)
类别详情
| ID | 是否混合 | 语音状态 | 环境状态 | 类别标签 | 描述 |
|---|---|---|---|---|---|
| 0 | ❌ | 真实 | 真实 | original | 未经混合的原始真实语音及对应环境音频 |
| 1 | ✅ | 真实 | 真实 | bonafide_bonafide | 真实语音与另一真实环境音频混合 |
| 2 | ✅ | 欺骗 | 真实 | spoof_bonafide | 欺骗语音与真实环境音频混合 |
| 3 | ✅ | 真实 | 欺骗 | bonafide_spoof | 真实语音与欺骗环境音频混合 |
| 4 | ✅ | 欺骗 | 欺骗 | spoof_spoof | 欺骗语音与欺骗环境音频混合 |
数据来源
- 真实语音: ASVspoof5、CommonVoice
- 欺骗语音: ASVspoof5、SSTC
- 真实环境声音: VGGSound
- 欺骗环境声音: VCapAV
- 原始混合音频: VGGSound(同时捕获语音和环境)
环境声音涵盖室内、街道和自然环境,确保声学多样性。
数据处理
- 所有文件重采样至16 kHz。
- 最终时长由较短信号决定,较长信号被截断。
- 环境声音相对于语音按预定义信噪比进行缩放。
下载信息
数据集下载链接:https://xuepingzhang.github.io/CompSpoof-dataset/

- 1CompSpoof: A Dataset and Joint Learning Framework for Component-Level Audio Anti-spoofing Countermeasures美国洛杉矶的OfSpectrum公司和中国昆山杜克大学数字创新研究中心 · 2025年



