遇见数据集

Tut Sound Events 2018 - Ambisonic, Reverberant And Synthetic Impulse Response Dataset

收藏
Zenodo2020-09-20 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>Tampere University of Technology (TUT) Sound Events 2018 - Ambisonic, Reverberant and Synthetic Impulse Response Dataset</strong> This dataset consists of simulated reverberant first order Ambisonic (FOA) format recordings with stationary point sources each associated with a spatial coordinate. The dataset consists of three sub-datasets with a) maximum one temporally overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240 recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), spatial location in azimuth and elevation angles (in degrees), and distance from the microphone (in meters). The sound events are spatially placed within a room using the image source method. The room size chosen was 10x8x4 meter with reverberation time per octave band of [1.0, 0.8, 0.7, 0.6, 0.5, 0.4] s and 125 Hz–4 kHz band center frequencies. The isolated sound events were taken from the DCASE 2016 task 2 dataset. This dataset consists of 11 sound event classes such as Clearing throat, Coughing, Door knock, Door slam, Drawer, Human laughter, Keyboard, Keys (put on a table), Page turning, Phone ringing and Speech. The sound events are randomly placed in a spatial grid with 10-degree resolution in full azimuth and [-60 60) degree elevation angles. Additionally, the sound events are placed at a random distance of at least 1 meter away from the microphone. The license of the dataset can be found in the LICENSE file. The rest of the nine zip files consists of datasets for a given split and overlap. For example, the ov3_split1.zip file consists of the audio and metadata folders for the case of maximum three temporally overlapping sound events (ov3) and the first cross-validation split (split1). Within each audio/metadata folder, the filenames for training split have the 'train' prefix, while the testing split filenames have the 'test' prefix. This dataset was collected as part of the 'Sound event localization and detection of overlapping sources using convolutional recurrent neural network' work.

<strong>坦佩雷理工大学(Tampere University of Technology, TUT)2018年声音事件数据集——第一阶环绕声(First Order Ambisonic, FOA)、混响与合成冲激响应数据集</strong> 本数据集包含模拟生成的混响型第一阶环绕声(First Order Ambisonic, FOA)格式录音,声源为固定点声源,每个声源均关联空间坐标。数据集包含三个子数据集,分别为a)最多存在1个时间重叠的声音事件,b)最多存在2个时间重叠的声音事件,以及c)最多存在3个时间重叠的声音事件。每个子数据集均包含3组交叉验证划分:训练划分包含240段时长约30秒的录音,测试划分包含60段等长录音。每段录音对应一份同名元数据文件,其中包含声音事件名称、时间起始与结束时刻(单位:秒)、方位角与俯仰角空间位置(单位:度),以及距麦克风的距离(单位:米)。声音事件通过镜像源法(image source method)放置于室内空间中,所选房间尺寸为10×8×4米,各倍频程混响时间为[1.0, 0.8, 0.7, 0.6, 0.5, 0.4]秒,中心频率覆盖125 Hz至4 kHz频段。 该数据集的孤立声音事件取自DCASE 2016任务2数据集,包含11类声音事件,分别为清嗓、咳嗽、敲门声、摔门声、抽屉开关声、人类笑声、键盘敲击声、放钥匙至桌面声、翻页声、电话铃声与语音。声音事件随机放置于空间网格中,方位角覆盖全角且分辨率为10度,俯仰角范围为[-60, 60)度,同时声源与麦克风的最小距离为1米。 本数据集的授权协议可参见LICENSE文件。其余9个压缩包均为对应划分与重叠数的数据集。例如,ov3_split1.zip文件包含最多3个时间重叠声音事件(ov3)且为第1组交叉验证划分(split1)的音频与元数据文件夹。在每个音频/元数据文件夹中,训练划分的文件名带有“train”前缀,测试划分的文件名带有“test”前缀。 本数据集是“基于卷积循环神经网络的声音事件定位与重叠声源检测”研究工作的一部分。

提供机构:
Zenodo
创建时间:
2018-06-30
二维码
社区交流群
二维码
科研交流群
商业服务