nEMO Polish Speech Emotion Dataset: 107 Acoustic Features
收藏资源简介:
This dataset accompanies the article “Emotion Recognition Using Acoustic Features and Deep Learning: A Speaker-Independent Study” by Marcin Kołodziej, Andrzej Majkowski, and Tomasz Rywik. The dataset contains 107 acoustic features extracted from 2327 Polish speech utterances from the nEMO corpus. The recordings represent three affective classes: sad, neutral, and happiness. The data were prepared for research on speech emotion recognition under speaker-independent evaluation conditions, especially leave-one-subject-out (LOSO) validation. Each row in the main CSV file corresponds to one utterance. The dataset includes speaker identifiers, utterance identifiers, emotion labels, class IDs, source file names, and 107 numeric acoustic descriptors. The acoustic features cover prosodic, phonatory, temporal, energy-related, spectral, and cepstral properties of speech, including F0 statistics, jitter, shimmer, harmonic-to-noise ratio, pause and onset measures, RMS energy, zero crossing rate, spectral centroid, bandwidth, rolloff, flatness, MFCCs, and delta MFCCs. The release is intended to support reproducibility of the feature-based analyses reported in the article and to enable further research on Polish speech emotion recognition, speaker-independent affect classification, acoustic feature analysis, and comparison of classical acoustic descriptors with deep learning approaches such as wav2vec 2.0 and WavLM. The package includes:- a main analysis-ready CSV table with all utterances and all 107 features,- a feature dictionary describing feature names, groups, descriptions, and units,- metadata in JSON format,- speaker-level CSV files,- utterance-level JSON files. Original absolute filesystem paths were removed from the public release. The dataset contains acoustic feature values and metadata only; it does not include the original audio recordings.



