该数据集名为ASR Leaderboard: Longform Test Sets,包含三个长格式自动语音识别(ASR)基准测试集:Earnings-21、Earnings-22和TED-LIUM。这些数据集用于评估在现实条件下(如长时间音频片段、重叠说话人和特定领域语言)的长格式ASR模型性能。数据集以标准化的Parquet格式提供,包含音频、文本以及根据数据集不同而异的额外元数据。README
This resource is available for download in Kielipankki - The Language Bank of Finland as part of "Donate Speech: Selected dataset", http://urn.fi/urn:nbn:fi:lb-2022060127. The resource contains a subs
# Dataset Card for "AutomaticSpeechRecognition_LJSpeech" [More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)