遇见数据集

saurabh1043/noisy_kannada

收藏
Hugging Face2026-05-15 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个音频-文本对数据集,包含id、文本和音频三个特征字段。数据分为训练集(train)和开发集(dev),其中训练集有592,047个样本,总大小约163GB;开发集有31,160个样本,总大小约9.18GB。整个数据集下载大小约为171GB,总数据集大小约为172GB。

This dataset is an audio-text pair dataset containing three feature fields: id, text, and audio. The data is divided into a training set (train) and a development set (dev), with the training set consisting of 592,047 samples and a total size of approximately 163GB, and the development set consisting of 31,160 samples and a total size of approximately 9.18GB. The entire dataset has a download size of about 171GB and a total dataset size of about 172GB.

提供机构:
saurabh1043
二维码
社区交流群
二维码
科研交流群
商业服务