The SWARA Speech Corpus: A Large Parallel Romanian Read Speech Dataset
收藏资源简介:
Please be aware that ACCESS will be granted only if you clearly state your affiliation (including link to personal homepage or LinkedIn Account) and end-use of the corpus! The SWARA Corpus is a result of the SWARA Project, funded by the Romanian Ministry of Education, under the grant agreement PN-II-PT-PCCA-2013-4 No 6/2014. The corpus contains over 21 hours of high quality recordings from 17 different speakers. The data is segmented in 19,279 utterances and includes their orthographic transcripts and semi-automatic phone-level alignments. If you use the SWARA Corpus, please cite the following paper: Adriana Stan, Florina Dinescu, Cristina Țiple, Șerban Meza, Bogdan Orza, Magdalena Chirilă and Mircea Giurgiu, The SWARA Speech Corpus: A Large Parallel Romanian Read Speech Dataset, in Proceedings of the 9th Conference on Speech Technology and Human-Computer Dialogue, Bucharest, Romania, July 6-9, 2017 pdf | bib You can listen to audio samples of each speaker, as well as samples of synthetic voices built from the SWARA corpus HERE



