遇见数据集

Aliasing vs Non-Aliasing Audio Dataset

收藏
Zenodo2025-10-09 更新2026-05-26 收录
官方服务:

资源简介:

1. Source Data & Chunking Datasets: UrbanSound8K (10 folds of short environmental clips), ESC-50 (50 classes of environmental sounds), and a MAVD city-recording corpus. J. Salamon, C. Jacoby and J. P. Bello, "A Dataset and Taxonomy for Urban Sound Research", 22nd ACM International Conference on Multimedia, Orlando USA, Nov. 2014. K. J. Piczak. ESC: Dataset for Environmental Sound Classification. Proceedings of the 23rd Annual ACM Conference on Multimedia, Brisbane, Australia, 2015. Pablo Zinemanas, Pablo Cancela and Martín Rocamora. MAVD: a dataset for sound event detection in urban environments. DCASE 2019 Workshop, 25-26 October 2019, New York, USA Chunk duration: Each audio file is broken into 1 s segments (using only one channel of stereo). Aliasing selection: Chunks whose high-frequency energy (>11 kHz) exceeds a small fraction of total energy are flagged as aliasing present; the rest are discarded. 2. Filtered vs. Unfiltered Unfiltered: Raw 1 s chunks with aliasing content retained. Filtered: Same chunks passed through a 6-pole Butterworth low-pass at Nyquist (sample rate/2), removing aliasing content. 3. Multi-Rate Resampling For each selected chunk (filtered & unfiltered), you resampled to six different rates: 1000 Hz, 2000 Hz, 4000 Hz, 8000 Hz, 16000 Hz, 22000 Hz This lets you study the impact of down-sampling on both aliased and alias-free audio. 4. Feature Extraction You built two parallel feature pipelines: STFT Parameters: 512-point FFT, 50% overlap. Saved the magnitude spectrograms (2D arrays of shape [freq_bins × time_frames]). Two-pass normalization (global min/max) ensures every spectrogram is scaled identically. MFCC Parameters: 20 coefficients, FFT window = sr/2, hop = sr/4. Flattened the 20 × time_frames MFCC matrix to 1D. Z-score normalized per coefficient (using training split stats) for consistency. 5. Train / Validation / Test Splits 80 % / 10 % / 10 % split of the selected chunks, stratified by sample type (filtered vs. unfiltered) and class. Directory layout (example for STFT at 8 kHz): Processed_Files/ DS_U8K/ # also ESC50 and Zen 8000/ # and all other sample rates train/ filtered/ ← .npy spectrograms unfiltered/ validation/ filtered/ unfiltered/ test/ filtered/ unfiltered/ Each file named <chunk_index>-<class>.npy, where <class> is the original label extracted from the filename. The MAVD city-sound recording corpus does not have a class as the original dataset was not class-based.

提供机构:
Zenodo
创建时间:
2025-08-02
二维码
社区交流群
二维码
科研交流群
商业服务