atiyehghm/ganjoor_stt_overlap
收藏官方服务:
资源简介:
该数据集是一个用于语音识别或音频处理任务的数据集,包含音频文件及其对应的文本转录。特征包括音频数据、转录文本、单词计数、字符计数、音频时长和采样率。数据集分为训练集和测试集,训练集有102,573个示例,测试集有43,961个示例,总大小约为55.7 GB。
This dataset is designed for speech recognition or audio processing tasks, containing audio files along with their corresponding text transcripts. Features include audio data, transcript text, word count, character count, audio duration, and sample rate. The dataset is split into training and test sets, with 102,573 examples in the training set and 43,961 examples in the test set, totaling approximately 55.7 GB in size.
提供机构:
atiyehghm


