saurabh1043/noisy_kannada
收藏官方服务:
资源简介:
该数据集是一个音频-文本对数据集,包含id、文本和音频三个特征字段。数据分为训练集(train)和开发集(dev),其中训练集有592,047个样本,总大小约163GB;开发集有31,160个样本,总大小约9.18GB。整个数据集下载大小约为171GB,总数据集大小约为172GB。
This dataset is an audio-text pair dataset containing three feature fields: id, text, and audio. The data is divided into a training set (train) and a development set (dev), with the training set consisting of 592,047 samples and a total size of approximately 163GB, and the development set consisting of 31,160 samples and a total size of approximately 9.18GB. The entire dataset has a download size of about 171GB and a total dataset size of about 172GB.
提供机构:
saurabh1043


