hf-audio/open-asr-leaderboard
收藏资源简介:
该数据集是ESB(开放ASR排行榜)测试集的排序和Parquet格式版本,源自open-asr-leaderboard/datasets-test-only数据集。它包含九个配置:ami、common_voice、earnings22、gigaspeech、librispeech、spgispeech、tedlium、voxpopuli和voxpopuli_cleaned_aa,每个配置都按音频长度排序,并将数据格式从自定义加载脚本转换为安全的Parquet格式。数据集用于自动语音识别(ASR)系统的评估,每个数据点包括音频(采样率16kHz)、转录文本、唯一ID、数据集名称和音频长度(秒)。音频已分段适合训练,转录经过纠错处理(如去除垃圾标记、转换标点符号)。测试集的转录未提供,需通过Hugging Face空间提交预测结果进行评分。所有数据集可自由访问,但Common Voice、GigaSpeech和SPGISpeech需同意特定使用条款。数据集还附带一个8小时的诊断数据集,用于跨域验证。ESB数据集涵盖多种领域(如有声书、会议、TED演讲)和说话风格(如叙述、即兴),总训练时长从78小时到4900小时不等。
This dataset is a sorted, Parquet-formatted version of the ESB (Open ASR Leaderboard) test set, derived from the open-asr-leaderboard/datasets-test-only dataset. It consists of nine configurations: ami, common_voice, earnings22, gigaspeech, librispeech, spgispeech, tedlium, voxpopuli, and voxpopuli_cleaned_aa. Each configuration is sorted by audio length, with the data format converted from custom loading scripts to secure Parquet format. This dataset is designed for evaluating automatic speech recognition (ASR) systems. Each data point includes audio with a 16kHz sampling rate, transcript text, unique ID, dataset name, and audio length measured in seconds. Audio segments are optimized for training, and transcripts have been post-processed with corrections such as removing garbage markers and standardizing punctuation marks. Transcripts for the test set are not publicly provided; predictions must be submitted via Hugging Face Spaces for official scoring. All datasets are freely accessible, but Common Voice, GigaSpeech, and SPGISpeech require users to agree to specific terms of use. Additionally, the dataset package includes an 8-hour diagnostic dataset for cross-domain validation. The ESB dataset covers diverse domains including audiobooks, meetings, TED talks, and various speaking styles such as narration and impromptu speech, with total training durations ranging from 78 hours to 4900 hours across different configurations.




