CL-MASR
收藏资源简介:
<strong>CL-MASR Dataset</strong> This is the dataset used in the continual learning for multilingual ASR (CL-MASR) benchmark. It is composed of speech recordings from 20 languages selected from the Common Voice 13 dataset. For each language, it includes up to 10/1/1 hours for train/dev/test, respectively. The CL-MASR benchmark platform is available in the SpeechBrain toolkit (see recipes/CommonVoice):<br> https://github.com/speechbrain/speechbrain The original Common Voice 13 data are available at:<br> https://commonvoice.mozilla.org/en/datasets <strong>List of Languages</strong> - English (en)<br> - Chinese (zh-CN)<br> - German (de)<br> - Spanish (es)<br> - Russian (ru)<br> - French (fr)<br> - Portuguese (pt)<br> - Japanese (ja)<br> - Turkish (tr)<br> - Polish (pl)<br> - Kinyarwanda (rw)<br> - Esperanto (eo)<br> - Kabyle (kab)<br> - Luganda (lg)<br> - Meadow Mari (mhr)<br> - Central Kurdish (ckb)<br> - Abkhaz (ab)<br> - Kurmanji Kurdish (kmr)<br> - Frisian (fy-NL)<br> - Interlingua (ia)



