LHCP-ASR
收藏资源简介:
该数据集包含两种配置:'longform'和'segments',均包含音频及其转录文本。'longform'配置的训练集包含560个样本,总大小约48.5GB;开发集和测试集分别包含2020年和2022年的子集,样本数量从11到32不等。'segments'配置的训练集包含174,376个样本,总大小约40GB;开发集和测试集同样分为2020年和2022年子集,样本数量从3,622到9,738不等。数据集适用于语音识别、语音合成等音频处理任务。
This dataset includes two configurations: 'longform' and 'segments', both comprising audio recordings and their corresponding transcriptions. For the 'longform' configuration, the training set consists of 560 samples with a total size of approximately 48.5 GB. The development and test sets are respectively subsets from 2020 and 2022, with the number of samples ranging from 11 to 32. For the 'segments' configuration, the training set contains 174,376 samples with a total size of around 40 GB. The development and test sets are also split into 2020 and 2022 subsets, with sample counts varying from 3,622 to 9,738. This dataset is suitable for audio processing tasks such as speech recognition and speech synthesis.




