遇见数据集

huggingworld/open-asr-leaderboard

收藏
Hugging Face2026-04-26 更新2026-05-03 收录
官方服务:

资源简介:

该数据集名为ESB Test Sets: Parquet & Sorted,是一个用于自动语音识别(ASR)评估的测试集集合,基于open-asr-leaderboard/datasets-test-only数据集生成,主要特点是将数据按音频长度排序并转换为parquet格式以提高安全性和易用性。数据集包含8个子数据集:AMI(会议录音)、Common Voice(维基百科朗读)、Earnings22(财报电话会议)、GigaSpeech(有声书、播客和YouTube内容)、LibriSpeech(有声书)、SPGISpeech(财务会议)、TED-LIUM(TED演讲)和VoxPopuli(欧洲议会演讲),每个子数据集提供测试集,数据字段包括音频(采样率16kHz)、文本转录、数据集名称和唯一ID。数据集旨在简化ASR系统的训练和评估,支持通过Hugging Face Datasets库一键下载和准备,部分子数据集需要特定使用协议。此外,数据集还包含一个诊断数据集,用于跨领域评估。所有数据均经过错误校正处理,转录文本未提供测试集以支持公平评估。

This dataset is named ESB Test Sets: Parquet & Sorted, a collection of test sets for automatic speech recognition (ASR) evaluation, derived from the open-asr-leaderboard/datasets-test-only dataset. It is characterized by sorting each split by audio length and converting the format to parquet for improved safety and usability. The dataset includes eight sub-datasets: AMI (meeting recordings), Common Voice (Wikipedia narration), Earnings22 (earnings calls), GigaSpeech (audiobooks, podcasts, and YouTube content), LibriSpeech (audiobooks), SPGISpeech (financial meetings), TED-LIUM (TED talks), and VoxPopuli (European Parliament speeches). Each sub-dataset provides a test split with features such as audio (sampling rate 16kHz), text transcription, dataset name, and unique ID. It is designed to streamline training and evaluation of ASR systems, allowing one-line download and preparation via the Hugging Face Datasets library, with some sub-datasets requiring specific usage agreements. Additionally, the dataset includes a diagnostic set for cross-domain evaluation. All data is error-corrected, and transcriptions are not provided for test splits to ensure fair assessment.

提供机构:
huggingworld
二维码
社区交流群
二维码
科研交流群
商业服务