遇见数据集

reddyrohith49471/jt-dataset-final3

收藏
Hugging Face2026-04-29 更新2026-05-03 收录
官方服务:

资源简介:

该数据集是一个多语言语音数据集,包含音频文件及其对应的文本转录。每个样本由以下特征组成:音频数据(采样率为16000Hz)、对应的句子文本、说话人ID标识以及语言标签。数据集分为训练集(896个样本)和测试集(100个样本),总大小约为680MB,适用于语音识别、说话人识别或多语言语音处理任务。

This dataset is a multilingual speech dataset containing audio files and their corresponding text transcriptions. Each sample consists of the following features: audio data (with a sampling rate of 16000 Hz), corresponding sentence text, speaker ID, and language label. The dataset is divided into a training set (896 examples) and a test set (100 examples), with a total size of approximately 680MB, suitable for speech recognition, speaker identification, or multilingual speech processing tasks.

提供机构:
reddyrohith49471
二维码
社区交流群
二维码
科研交流群
商业服务