Speech-to-text dataset for keyword spotting applications in the Papiamento language within a healthcare environment
收藏资源简介:
This audio dataset accompanies the SISSTEM Bachelor's thesis "Speech-to-text model for keyword spotting applications in the Papiamento language within a healthcare environment" by Joel Rajnherc. It was developed to kick-start the exploration of keyword spotting research for the Papiamento language. The speech_commands dataset heavily inspired it. It currently consists of seven classes, six keywords, and one class for unknowns. Approximately 280 participants contributed to the dataset, which resulted in 16766 samples after filtering. What do the instances in this dataset represent? Each instance represents a one-second-long recording, stored as a spectrogram with 124 frequency bins by 129 time frames. Each sample is stored in the .npz file format. Regarding the target label, there are seven classes in total: Papiamento English Dolor Pain Masha danki Thank you very much No No Resultado Result SSIMSAN (designated wake word) SSIMSAN (designated wake word) Unknown (auxiliary bucket) Unknown (auxiliary bucket) Can I reproduce the results easily? Certainly, we provide a notebook code repository for educational and academic interests here. Acknowledgements All the participants who contributed to the data collection! All the stakeholders, including but not limited to Full Stack Vision Aruba (Raspberry PI Foundation), ImSan, FHTMS, and Esther Plomp from the University of Aruba Research Center. The thesis supervisor Francis Laclé, co-evaluators Dr. Salys Sultan, Dr. Eric Mijts, and external readers Peter Scholing and John Becker. A special thank you to Jean-Luc at SISSTEM for providing help with data filtering.



