Small VCTK subset
收藏资源简介:
This dataset is a very small subsample of the CSTR VCTK Corpus, originally published by the University of Edinburgh (reference included below). The motive for this subset is to enable a very small (~220MB compressed) dataset for easy prototyping and testing. The audio has been downsampled to 8kHz and only inputs from mic1 are kept (original dataset contains stereo audio labeled as mic1 and mic2). Furthermore, only audio files between 2 and 5 seconds were kept, and a random subset of 50 files per each speaker was selected. Properties of the small VCTK subset: Dataset is accompanied by a csv file containing the speaker ID, filename, transcript, gender, age, and accent for each file in the dataset. Filenames are consistent with original VCTK corpus. Predefined train, test, & validation splits (exclusive per-speaker, i.e. a speaker can only belong to one of the data split). All recordings from original VCTK were filtered and contain only mic1 data. All recordings from original VCTK were filtered and contain only audio of duration between 2 and 5 seconds. All audio was subsampled to 8kHz. Dataset contains 108 speakers as speakers p280 and p315 were removed from original dataset.Train split: 78 speakers, test split: 15 speakers, validation split: 15 speakers. Contains 50 randomly selected recordings per each speaker, i.e. 108x50 = 5400 recordings in total. Reference to original VCTK corpus: @misc{yamagishi2019vctk, author = {Yamagishi, Junichi and Veaux, Christophe and MacDonald, Kirsten}, title = {{CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit}}, year = {2019}, version = {0.92}, note = {doi: doi.org/10.7488/ds/2645}, institution = {University of Edinburgh, The Centre for Speech Technology Research (CSTR)}}



