遇见数据集

Punjabi Speech: A labeled Speech Corpus

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

The Punjabi Speech corpus is designed for automatic speech recognition and synthesis purposes. The corpus comprises recorded speech samples in the studio and open environment settings, with a sampling rate of 44.1 kHz in WAV file format. The duration of each recording is limited to 15 seconds to prevent memory issues while training on GPUs. The dataset currently contains 2429 spoken utterances from two male speakers, totaling ~4 hours of data. For training, validation, and testing purposes, the data is pre-divided into 80% for training, 10% for validation, and 10% for testing. The dataset is organized in a straightforward manner, with all speech files located in the "clips" directory and transcript files (train, dev, and test) in TSV format located in the parent directory. Each line in the transcript files represents a label for a single speech sample in the clips directory. The first column contains the path/name to the corresponding WAV file and the second column, separated by a tab, contains the transcript in text form.

本旁遮普语语音语料库(Punjabi Speech Corpus)专为自动语音识别与语音合成任务研发。该语料库包含演播室与开放环境下录制的语音样本,所有音频均采用WAV文件格式,采样率为44.1 kHz。为避免GPU训练时出现内存溢出问题,单条录音时长被限制为15秒。当前数据集包含两名男性发音者的2429条口语话语,总时长约4小时。针对训练、验证与测试需求,该数据集已预先划分为训练集(占比80%)、验证集(占比10%)与测试集(占比10%)。该数据集组织方式简洁直观:所有语音文件均存放于"clips"目录下,训练、验证与测试对应的转写文件均采用TSV格式,存放在上级目录中。转写文件中的每一行对应"clips"目录下的一条语音样本的标注:第一列为对应WAV文件的路径/文件名,第二列以制表符分隔,存储该语音的文本转写内容。

创建时间:
2023-08-16
二维码
社区交流群
二维码
科研交流群
商业服务