CSS10 is a collection of single speaker speech datasets for 10 languages. Each of them consists of audio files recorded