遇见数据集

augCENSE-18k

收藏
Zenodo2021-06-01 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

Created By Félix Gontier and Mathieu Lagrange, LS2N, CNRS, Ecole Centrale Nantes Contact : mathieu.lagrange@cnrs.fr If used for research, please refer to: <pre>@article{gontier2021training, title={Polyphonic training set synthesis improves self-supervised urban sound classification}, author={Félix Gontier and Vincent Lostanlen, and Mathieu Lagrange and Nicolas Fortin and Jean-Francois Petiot and Catherine Lavandier}, journal={The Journal of the Acoustical Society of America}, year={2021}, publisher={Acoustical Society of America} } </pre> augCENSE-18k is a derivative of CENSE-2k, obtained by time stretching and pitch shifting audio clips of the \emph{voice} and \emph{birds} classes at random.<br> The total duration of the dataset is equal to 18k seconds, i.e., the same as simCENSE-18k, with balanced material over classes. Each audio samples are cut into one or several 3 seconds parts, each resulting into spectrograms of size 23x29, leading to a dataset of 609 spectrograms. Low volume amorphic background noise recordings is added and the cut audio sample is centered within the 3 seconds if shorter. &gt;&gt;&gt; a=numpy.load('augCENSE-18k_train_spectralData.npy') &gt;&gt;&gt; a.shape (4421, 23, 29) &gt;&gt;&gt; a=numpy.load('augCENSE-18k_train_presence.npy') &gt;&gt;&gt; a.shape (4421, 16, 3) The 3 dimensions corresponds to the sceneId, the frameId (time), the sourceId (traffic, voice, birds). Annotation is provided as a binary indicator of source presence for one second, that is 8 consecutive 125 ms frames with a hop of one frame.

本数据集由费利克斯·贡捷(Félix Gontier)与马蒂厄·拉格朗日(Mathieu Lagrange)创建,所属机构为LS2N实验室、法国国家科学研究中心(CNRS)、南特中央理工学院(École Centrale Nantes)。联系方式:mathieu.lagrange@cnrs.fr。若用于学术研究,请引用如下文献: <pre>@article{gontier2021training, title={Polyphonic training set synthesis improves self-supervised urban sound classification}, title={多声部训练集合成优化自监督城市声分类}, author={Félix Gontier and Vincent Lostanlen, and Mathieu Lagrange and Nicolas Fortin and Jean-Francois Petiot and Catherine Lavandier}, author={费利克斯·贡捷、樊尚·洛斯坦伦、马蒂厄·拉格朗日、尼古拉·福坦、让-弗朗索瓦·佩蒂奥、卡特琳·拉凡迪耶}, journal={The Journal of the Acoustical Society of America}, journal={《美国声学学会杂志》}, year={2021}, publisher={Acoustical Society of America} publisher={美国声学学会} }</pre> augCENSE-18k是CENSE-2k的衍生数据集,通过对**语音(voice)**与**鸟鸣(birds)**类别的音频片段进行随机时域拉伸与音调偏移处理得到。 本数据集总时长为18000秒,与simCENSE-18k的总时长一致,且各类别样本分布均衡。将每条音频剪辑为一段或多段3秒时长的片段,由此生成尺寸为23×29的语谱图(spectrogram),最终得到由609张语谱图组成的数据集。 向数据集添加低幅度无定形背景噪声录音;若原始音频片段短于3秒,则将其居中放置于3秒时长的区间内。 >>> a=numpy.load('augCENSE-18k_train_spectralData.npy') >>> a.shape (4421, 23, 29) >>> a=numpy.load('augCENSE-18k_train_presence.npy') >>> a.shape (4421, 16, 3) 数组的三个维度分别对应:场景ID(sceneId)、帧ID(时间维度,frameId)、声源ID(sourceId,涵盖交通声traffic、语音声voice、鸟鸣声birds)。标注以二进制指示符形式给出声源在1秒时长内的存在状态,即包含8个连续的125毫秒帧,帧步长为1帧。

提供机构:
Zenodo
创建时间:
2021-06-01
二维码
社区交流群
二维码
科研交流群
商业服务