遇见数据集

<strong>CMU-synth: A synthesized Punjabi Speech dataset </strong>

收藏
DataCite Commons2023-06-30 更新2024-08-18 收录
官方服务:

资源简介:

The CMU-synth dataset is a synthetic Punjabi dataset generated using CMU's Clustergen text-to-speech model. It consists of approximately 80,000 synthesized utterances featuring a single synthetic female speaker, resulting in around 170 hours of audio data. The dataset has been partitioned into three sets for training, validation, and testing, with proportions of 80%, 10%, and 10% respectively. The dataset is well-organized, with all speech files stored in the "clips" directory. The corresponding transcript files for training, validation (dev), and testing are located in the parent directory, following the TSV (Tab-Separated Values) format. Each line in the transcript files represents a label assigned to a specific speech sample from the clips directory. The first column of each line contains the name of the corresponding WAV file, while the second column, separated by a tab, contains the transcript in textual format.

CMU-synth 数据集是依托CMU的Clustergen文本到语音(text-to-speech)模型生成的合成旁遮普语数据集。该数据集包含约8万条由单一合成女性发音人生成的合成语音片段,总音频时长约170小时。该数据集已被划分为训练集、验证集与测试集,三者占比分别为80%、10%与10%。该数据集组织规范,所有语音文件均存储于"clips"目录下。对应训练、验证(开发集,dev)与测试集的转录文件均位于父目录中,采用TSV(Tab-Separated Values,制表符分隔值)格式。转录文件的每一行对应"clips"目录下一条语音样本的标注信息,每行第一列为对应WAV(Waveform Audio File Format)文件的文件名,第二列以制表符分隔,为文本形式的转录内容。

提供机构:
figshare
创建时间:
2023-06-30
搜集汇总
数据集介绍
<strong>CMU-synth: A synthesized Punjabi Speech dataset </strong> 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务