Youtube-Dataset for Language Identification in Speech Signals
收藏资源简介:
<strong>Youtube-Dataset for Language Identification in Speech Signals</strong> - for scientific use only, for questions contact: jakob.abesser@idmt.fraunhofer.de <strong>Reference</strong> In case you use this dataset for your research, please cite Alexandra Draghici, Jakob Abeßer & Hanna Lukashevich: A Study on Spoken Language Identification<br> using Deep Neural Networks, Proceedings of the Audio Mostly Conference 2020 <strong>Dataset</strong> The YouTube News Collection is a collection of videos from various<br> Youtube news channels. We gathered data from channels like BBC<br> news, France24, DW News, and Noticias Telemundo. - 135664 npy files (numpy matrices exported from Python)<br> - each npy file includes a mel spectrogram (see below) of an audio file<br> - the subfolders "0" - "5" encode the language id:<br> 0 - English<br> 1 - French<br> 2 - German<br> 3 - Greek<br> 4 - Italian<br> 5 - Spanish <strong>Audio Processing</strong> - mono, sample rate 22.05 kHz<br> - mel spectrogram (librosa python package)<br> - windows size 512 samples<br> - hopsize 441 samples (20 ms)<br> - 129 mel bands<br> - file-level spectrogram are normalized to maximum of 1<br>



