SHAL
收藏资源简介:
SHAL数据集是由上海声学实验室创建,专注于中文数字字符串的长文本依赖语音验证。该数据集包含约72.3小时的音频,共46,583个文件,格式为44.1kHz、16位PCM-WAV。数据集主要关注10至40岁的说话者,性别平衡。创建过程中,使用了Tacotron2和HiFi-GAN进行数据增强,通过转移学习和个性化TTS模型,将数据集扩展至原大小的六倍。SHAL数据集适用于金融支付等领域的身份验证,旨在解决文本依赖语音验证中的数据稀缺和领域不匹配问题。
The SHAL Dataset was developed by the Shanghai Acoustic Laboratory, focusing on long-text-dependent speech verification for Chinese digital strings. This dataset contains approximately 72.3 hours of audio, totaling 46,583 files, with the format of 44.1kHz, 16-bit PCM-WAV. It mainly targets speakers aged 10 to 40, with a balanced gender distribution. During the dataset creation, data augmentation was conducted using Tacotron2 and HiFi-GAN, and the dataset was expanded to six times its original size via transfer learning and personalized TTS models. The SHAL Dataset is suitable for identity verification scenarios such as financial payment, aiming to address the problems of data scarcity and domain mismatch in text-dependent speech verification.

- 1A text-dependent speaker verification application framework based on Chinese numerical string corpus上海声学实验室,中国科学院,上海 · 2023年



