SingBAP: A Multimodal Dataset of Singing Performance with Biosignals, Audio, and Pose Estimation
收藏资源简介:
SingBAP is a multimodal dataset designed to support research in singing voice analysis, vocal pedagogy, and cross-modal machine learning. The dataset captures synchronised recordings of singing performance across multiple sensing modalities, enabling the study of relationships between audio output, posture, and physiological signals during vocal production. The dataset includes audio recordings captured from three devices (a 2017 MacBook Air, an iPhone 14 Pro, and a Behringer C-3 microphone), video recordings from two devices (a 2017 MacBook Air and an iPhone 14 Pro), as well as biosignal data including electromyography (EMG), electroencephalography (EEG) and piezoelectric respiration (PZT) signals. These modalities were recorded in a synchronised manner to facilitate multimodal alignment and analysis. Data were collected from 14 participants spanning professional, intermediate, and beginner singing experience levels. Participants performed structured vocal exercises, including scales with varying vowel sounds, note durations, and pitch ranges, providing variability in the output. The dataset is accompanied by a full processing pipeline implemented in Python and Jupyter Notebooks, including tools for data synchronisation, cleaning, feature extraction, and machine learning experimentation. There are three associated repositories with this dataset: vocal_data_feature_extraction audio_embeddings_and_feature_extraction_from_audio_dataset vocal_data_cleaning SingBAP is intended as a resource for research in multimodal machine learning, singing voice analysis, and educational technology development, particularly in applications that aim to model or support vocal technique learning and feedback systems.



