nvidia/hifitts-2
收藏资源简介:
HiFiTTS-2是一个大规模的高带宽语音数据集,从LibriVox有声读物中衍生而来。该数据集包含了大约36.7k小时的音频,来自5000名演讲者,音频可以从LibriVox以48 kHz的采样率下载。数据集的元数据包含了估计的带宽,用于推断录音的原始采样率。基础数据集经过过滤,适用于22 kHz的语音模型训练,同时也提供了一个适用于44 kHz训练的预计算子集。用户可以修改下载脚本来使用任何采样率和带宽阈值,这可能更适合他们的工作。
HiFiTTS-2 is a large-scale high bandwidth speech dataset derived from LibriVox audiobooks. The dataset contains approximately 36.7k hours of audio from 5k speakers that can be downloaded from LibriVox at a 48 kHz sampling rate. The metadata includes an estimated bandwidth, which is used to infer the original sampling rate the audio was recorded at. The base dataset is filtered for a bandwidth appropriate for training speech models at 22 kHz, and a precomputed subset is provided for 44 kHz training. Users can modify the download script to use any sampling rate and bandwidth threshold that might be more appropriate for their work.



