<strong>CMU-synth: A synthesized Punjabi Speech dataset </strong>
收藏资源简介:
The CMU-synth dataset is a synthetic Punjabi dataset generated using CMU's Clustergen text-to-speech model. It consists of approximately 80,000 synthesized utterances featuring a single synthetic female speaker, resulting in around 170 hours of audio data. The dataset has been partitioned into three sets for training, validation, and testing, with proportions of 80%, 10%, and 10% respectively. The dataset is well-organized, with all speech files stored in the "clips" directory. The corresponding transcript files for training, validation (dev), and testing are located in the parent directory, following the TSV (Tab-Separated Values) format. Each line in the transcript files represents a label assigned to a specific speech sample from the clips directory. The first column of each line contains the name of the corresponding WAV file, while the second column, separated by a tab, contains the transcript in textual format.



