jkot/parliament_hearings_processed
收藏官方服务:
资源简介:
该数据集是一个预处理过的议会听证会自动语音识别(ASR)数据集,包含音频和转录文本。音频的采样率为16000Hz。数据集分为训练集和测试集,训练集包含191455个样本,测试集包含2726个样本。数据集的下载大小为51507735963字节,数据集大小为51997848307字节。
This dataset is a preprocessed parliamentary hearing automatic speech recognition (ASR) dataset containing audio and transcribed text. The audio has a sampling rate of 16000 Hz. The dataset is split into training and test sets, with the training set including 191455 samples and the test set including 2726 samples. The download size of the dataset is 51507735963 bytes, and the total storage size of the dataset is 51997848307 bytes.
提供机构:
jkot原始信息汇总
数据集概述
数据集特征
- id: 数据类型为字符串。
- audio: 数据类型为音频,采样率为16000 Hz。
- transcription: 数据类型为字符串。
数据集分割
- 训练集 (train):
- 示例数量: 191455
- 数据大小: 53645064353.18 字节
- 测试集 (test):
- 示例数量: 2726
- 数据大小: 740331298.0 字节
数据集大小
- 下载大小: 51507379112 字节
- 数据集总大小: 54385395651.18 字节



