遇见数据集

SecureVoice50: A Multi-Device, Multi-Environment Speech Dataset for Robust and Secure Speaker Verification

收藏
Zenodo2025-09-30 更新2026-05-26 收录
官方服务:

资源简介:

This dataset (“SecureVoice50”) is intended to support research in speaker identification and voice biometric verification with an emphasis on security-sensitive applications (e.g., mobile banking). The dataset comprises approximately one hour of speech per participant, from 50 participants (20 female, 30 male; ages 21–38; diverse linguistic/cultural backgrounds) with varying English proficiency (advanced, intermediate, basic). Voice recordings were made using multiple devices: a low-end feature phone, two mid-range Android smartphones, iPhone SE, external USB condenser microphone with MacBook, and in-call capture on the feature phone to represent channel/network distortions. The speech was recorded in both indoor (semi-controlled lobby) and outdoor (traffic, wind, ambient noise) environments to reflect real-world conditions. Each recording is a .wav file (8–16 kHz sampling rates) labeled by speaker, device, environment, and utterance number. A metadata file is provided (CSV/Excel) mapping anonymized speaker IDs to non-identifying attributes: gender, age, nationality, and English proficiency. Participants provided informed consent; no personally identifying information is included; participants may withdraw their data. Potential use cases include baseline speaker verification, degradation-aware evaluation, spoofing/attack simulation, liveness detection, and adaptation studies. Related publication:“Speaker identification for low-end devices: A secure voice biometric solution for mobile banking,” Oyewale O. O., Taenaka Y., Kadobayashi Y., to appear in Proceedings of the 2025 IEEE International Conference on Privacy Computing and Data Security (PCDS), Hakodate, Japan. Ethics / Consent Statement:All participants gave informed consent. Only non-sensitive demographic attributes were collected. Recordings are anonymized. No personal identification information is provided. Usage is intended for academic research. Validation / Quality Assurance:Recordings were manually checked for completeness, clipped audio, interruptions or excessive noise; signal and SNR analyses ensured minimum thresholds. In system-level tests, speaker identification using embeddings (Whisper, ECAPA) achieved high accuracy, precision, recall, F1-score (≈ 99.5%) and loss decreased from ~0.1654 to ~0.0412 over training.

提供机构:
Zenodo
创建时间:
2025-09-30
二维码
社区交流群
二维码
科研交流群
商业服务