遇见数据集

Third-octave band spectra and spectrograms of DataSEC sounds

收藏
Zenodo2025-10-17 更新2026-05-29 收录
官方服务:

资源简介:

This dataset supports the spectral analysis presented in the associated study and is based on the DataSEC dataset (DOI: 10.5281/zenodo.15393250), a previously developed and validated collection by the authors. DataSEC contains 5,048 .wav audio recordings covering 22 environmental sound classes, further subdivided into a total of 40 sub-classes. For each recording, both the third-octave band spectrum and spectrogram are provided to support detailed spectral exploration. To offer practical spectral references for each sound class, third-octave band spectra and spectrograms were computed for all recordings. Each sub-class is treated as an independent category. Signal processing was performed using MATLAB R2024b Update 5, primarily leveraging functions from the Signal Processing Toolbox. The frequency range was set between 20 Hz and 20 kHz, consistent with standard practices in environmental acoustics and the typical range of human hearing. Given the variability in sound levels across the recordings, absolute sound pressure levels (SPL) were not used directly. Instead, each third-octave spectrum was min-max normalized, highlighting relative spectral patterns across different sound types and enabling meaningful comparisons. To capture temporal spectral variations, third-octave band spectrograms were computed using a 50 ms Hann window with 50% overlap. Each spectrogram displays time on the x-axis (in seconds), third-octave center frequencies on the left y-axis, and unweighted SPL on the color scale. This representation allows for the identification of both transient and continuous sound events and facilitates the interpretation of spectral dynamics across time and frequency. The dataset is organized into two main folders: "Third_octave_band_spectra.zip" contains one spectrum per audio file, grouped by sound class "Third_octave_band_spectrogram.zip" containing the corresponding spectrograms. All files are provided in .svg vector format to ensure maximum clarity and scalability. File names directly reference the corresponding .wav recordings in the original DataSEC dataset, ensuring seamless cross-referencing and traceability. The newer version contains updated figures in terms of visibility and readability.

本数据集支撑关联研究中的频谱分析工作,其基于作者团队此前开发并验证的DataSEC数据集(DOI: 10.5281/zenodo.15393250)。DataSEC包含5048条.wav音频录音,涵盖22类环境声音,进一步细分为总计40个子类别。针对每条录音,均提供了三分之一倍频带频谱(third-octave band spectrum)与语谱图(spectrogram),以支持精细化的频谱探索。 为给每一类声音提供实用的频谱参考,研究团队为所有录音计算了三分之一倍频带频谱与语谱图,且将每个子类别视为独立分类单元。 信号处理基于MATLAB R2024b Update 5完成,主要调用了信号处理工具箱(Signal Processing Toolbox)的相关函数。频谱频率范围设置为20 Hz至20 kHz,符合环境声学领域的通用标准与人类听觉的典型频段范围。 考虑到各录音间声级存在差异,研究未直接采用绝对声压级(SPL),而是对每条三分之一倍频带频谱进行了最小-最大归一化,以凸显不同声音类型间的相对频谱特征,实现更具意义的跨类别对比。 为捕捉时域频谱变化,研究采用50 ms汉宁窗(Hann window)、50%重叠率计算三分之一倍频带语谱图。每张语谱图以横轴表示时间(单位:秒),左侧纵轴标注三分之一倍频带中心频率,色阶则代表未加权声压级。该可视化方式可同时识别瞬态与持续声音事件,便于解读跨时域与频域的频谱动态特征。 本数据集分为两个核心文件夹: 1. "Third_octave_band_spectra.zip":按声音类别分组,每条音频文件对应一条三分之一倍频带频谱文件。 2. "Third_octave_band_spectrogram.zip":包含对应的语谱图文件。 所有文件均采用.svg矢量格式存储,以保障最高清晰度与可缩放性。文件名直接关联原始DataSEC数据集中对应的.wav录音文件,可实现无缝交叉引用与溯源。 本数据集的更新版本优化了可视化图形的辨识度与可读性。

提供机构:
Zenodo
创建时间:
2025-08-06
二维码
社区交流群
二维码
科研交流群
商业服务