TED-LIUM 3, Gigapeech, VoxPopuli-en
收藏资源简介:
本研究涉及三个主要的长格式英语自动语音识别(ASR)数据集:TED-LIUM 3、Gigapeech和VoxPopuli-en。这些数据集由Rev.com创建,旨在支持长格式ASR的研究,特别是解决训练与测试数据分割不一致的问题。数据集包含完整的录音和相应的转录,尽管转录的完整性有所不同。创建过程中,通过链接和扩展技术对这些数据集进行了重组,以增强其用于长格式ASR的能力。这些数据集的应用领域包括提高ASR模型在实际未分割音频上的性能,解决模型训练与实际应用中的不匹配问题。
This study involves three primary long-form English automatic speech recognition (ASR) datasets: TED-LIUM 3, Gigapeech, and VoxPopuli-en. These datasets were developed by Rev.com to support long-form ASR research, specifically targeting the resolution of inconsistent segmentation between training and test data. The datasets include complete audio recordings and their corresponding transcriptions, though the completeness of the transcriptions varies across samples. During their development, these datasets were restructured via linking and expansion techniques to improve their suitability for long-form ASR applications. The use cases of these datasets encompass enhancing the performance of ASR models on real-world unsegmented audio, as well as addressing the mismatch between model training and practical deployment.



