遇见数据集

saurabh1043/noisy_hindi

收藏
Hugging Face2026-05-15 更新2026-05-31 收录
官方服务:

资源简介:

这是一个大规模文本-音频配对数据集,包含590,236个训练样本和31,065个开发样本,总大小约134.5GB。每个样本包括id、文本内容和对应的音频数据,适用于语音识别、音频文本对齐等自然语言处理和音频处理任务。数据集以音频文件形式存储,支持机器学习模型的训练和评估。

This is a large-scale text-audio paired dataset containing 590,236 training samples and 31,065 development samples, with a total size of approximately 134.5 GB. Each sample includes an ID, text content, and corresponding audio data, suitable for natural language processing and audio processing tasks such as speech recognition and audio-text alignment. The dataset is stored in audio file format and supports the training and evaluation of machine learning models.

提供机构:
saurabh1043
二维码
社区交流群
二维码
科研交流群
商业服务