SHALCAS22A (简称SHAL)
收藏资源简介:
SHALCAS22A(简称SHAL)是由上海声学实验室创建的中文数值字符语料库,专为10至40岁年龄段的说话者设计,旨在支持金融交易中的说话者验证。该数据集包含约72.3小时的音频,共计46,583个文件,采用44.1kHz、16位PCM-WAV格式。创建过程中,选取了60名个体,每人提供25种不同文本类型的样本。SHAL数据集通过使用Tacotron2和HiFi-GAN进行数据增强,显著增加了数据多样性。该数据集主要应用于文本依赖型说话者验证(TD-SV),特别是在需要高精度等错误率(EER)表现的金融支付身份验证场景中。
SHALCAS22A (abbreviated as SHAL) is a Chinese numeric character corpus developed by the Shanghai Acoustic Laboratory, specifically designed for speakers aged 10 to 40 years old, with the goal of supporting speaker verification in financial transactions. This dataset contains approximately 72.3 hours of audio, totaling 46,583 files, formatted as 44.1kHz, 16-bit PCM-WAV. During its construction, 60 individual speakers were recruited, each providing samples for 25 distinct text types. The SHAL dataset adopts Tacotron2 and HiFi-GAN for data augmentation, which significantly improves the diversity of the dataset. It is primarily applied to text-dependent speaker verification (TD-SV), especially in financial payment identity verification scenarios that require high equal error rate (EER) performance.

- 1A framework of text-dependent speaker verification for chinese numerical string corpus上海声学实验室 · 2024年



