Mel Spectrogram Voice Dataset
收藏资源简介:
This dataset was obtained by applying a processing pipeline on two datasets Mozilla Common Voice and Cambridge IELTS Listening Audios. Mozilla Common Voice is a public, crowd sourced speech dataset collected from volunteers worldwide. It contains recordings in multiple languages, accents, and speaking styles, each paired with a text transcript. For this study, only English recordings were used, covering different accents, genders, and natural speaking variations. The original audio (48~kHz) was down sampled to 16~kHz to match the experimental setup. This dataset offers a diverse and natural collection of human speech. Cambridge IELTS recordings are studio-produced, scripted conversations and monologues with clear pronunciation and minimal background noise, making them similar to synthetic or machine-like voices in forensic analysis. All files were resampled to 16~kHz and split into 2-second segments for feature extraction. Together, these datasets provide a balanced mix of natural human speech and structured chatbot-like audio, enabling reliable testing of classification models. Furthermore, we apply deep learning models on MEL spectrogram data in order to detect if an audio segment is coming from a human or a chatbot.



