遇见数据集

A Multi-Condition Replay and Channel-Based Speech Spoofing Dataset Across 15 Bangladeshi Regions with Regional Accent Variability

收藏
DataONE2026-01-28 更新2026-02-07 收录
官方服务:

资源简介:

This dataset contains a large-scale speech spoofing corpus collected from speakers across 14 districts and the Sandwip Upazila of Bangladesh, capturing regional accent variability. The dataset includes both genuine and spoofed speech recorded under five real-world attack conditions: Monster, Monster–Telephone, Telephone, Robot, and Radio. Spoofed samples were generated through physical replay devices and communication channels, preserving realistic channel- and device-induced distortions. The audio files are organized into multiple subsets, including raw recordings, class-wise real vs. spoof partitions, and predefined training, validation, and testing splits. All recordings underwent quality control and preprocessing steps, including resampling, silence removal, and amplitude normalization. Due to file size constraints, the dataset is released as multiple compressed archives, each corresponding to specific spoofing scenarios and dataset splits. This dataset is intended to support research on speech anti-spoofing, replay attack detection, channel-aware speaker verification, and robustness evaluation of deep learning–based voice biometric systems.

本数据集为大规模语音欺骗语料库(speech spoofing corpus),采集自孟加拉国14个行政区及桑德威普乌帕齐拉(Sandwip Upazila)的多名说话人,覆盖区域口音的多样性。该数据集包含真实语音与欺骗语音两类样本,共涵盖五种真实世界攻击场景:Monster、Monster–Telephone、Telephone、Robot及Radio。欺骗样本通过物理回放设备与通信信道生成,完整保留了真实信道与设备引入的失真特征。音频文件被划分为多个子集,涵盖原始录音、按类别划分的真实/欺骗语音分区,以及预定义的训练集、验证集与测试集划分。所有录音均经过质量管控与预处理流程,包含重采样、静音去除与幅度归一化等步骤。受文件大小限制,本数据集以多个压缩归档包形式发布,每个归档包对应特定的欺骗攻击场景与数据集划分。本数据集旨在为语音反欺骗、回放攻击检测、信道感知型说话人验证,以及基于深度学习的语音生物识别系统的鲁棒性评估等相关研究提供支撑。

创建时间:
2026-01-30
二维码
社区交流群
二维码
科研交流群
商业服务