TAU Moving Sound Events 2019 - Ambisonic, Anechoic, Synthetic IR and Moving Source Dataset
收藏资源简介:
Tampere University (TAU) Moving Sound Events 2019 - Ambisonic, Anechoic and Synthetic Impulse Response (IR) and Moving Source Dataset This dataset consists of simulated anechoic first order Ambisonic (FOA) format recordings with moving point sources each in 2D spherical space represented with azimuth and elevation angles. The dataset consists of three sub-datasets with a) maximum one temporally overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240 recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), starting spatial location and directional spatial location in azimuth and elevation angles (in degrees), angular velocity of motion, and distance from the microphone (in meters). The isolated sound events were taken from the DCASE 2016 task 2 dataset. This dataset consists of 11 sound event classes such as Clearing throat, Coughing, Door knock, Door slam, Drawer, Human laughter, Keyboard, Keys (put on a table), Page turning, Phone ringing and Speech. Every event is assigned a spatial trajectory on an arc with a constant distance from the microphone (in the range 1-10 m) and moving with a constant angular velocity for its duration. Due to the choice of the ambisonic spatial recording format, the steering vectors for a plane wave source or point source in the far field are frequency-independent. Hence, there is no need for a time-variant convolution or impulse response interpolation scheme as the source is moving; the spatial encoding of the monophonic signal was done sample-by-sample using instantaneous ambisonic encoding vectors for the respective DOA of the moving source. The synthesized trajectories in the dataset vary in both azimuth and elevation and are simulated to have a constant angular velocity in the range [-90, 90]/s with 10-degree/s steps. The license of the dataset can be found in the LICENSE file. The rest of the nine zip files consists of datasets for a given split and overlap. For example, the ov3_split1.zip file consists of the audio and metadata folders for the case of maximum three temporally overlapping sound events (ov3) and the first cross-validation split (split1). Within each audio/metadata folder, the filenames for training split have the 'train' prefix, while the testing split filenames have the 'test' prefix. This dataset was collected as part of the 'Localization, Detection and Tracking of Multiple Moving Sound Sources with Convolutional Recurrent Neural Networks' work.
坦佩雷大学(Tampere University, TAU)2019年移动声事件数据集:Ambisonic、无回声与合成冲激响应(Impulse Response, IR)及移动源数据集 本数据集包含模拟生成的一阶Ambisonic(First Order Ambisonic, FOA)格式无回声录音,声源为二维球坐标系下的移动点声源,以方位角与俯仰角表征其空间位置。数据集包含三个子数据集,分别对应a) 最多1个时域重叠声事件、b) 最多2个时域重叠声事件,以及c) 最多3个时域重叠声事件。 每个子数据集均设置3组交叉验证划分,其中训练划分包含240段时长约30秒的录音,测试划分包含60段相同时长的录音。 每段录音配套同名元数据文件,文件内容包含声事件名称、时域起始与结束时间(单位:秒)、初始空间位置与终止空间位置的方位角和俯仰角(单位:度)、运动角速度,以及与麦克风的距离(单位:米)。 分离得到的声事件样本取自DCASE 2016任务2数据集。本数据集涵盖11类声事件,具体包括:清嗓、咳嗽、敲门声、摔门声、抽屉声、人类笑声、键盘声、钥匙放桌声、翻页声、电话铃声与语音。 每个声事件均被赋予一条空间运动轨迹:声源始终位于以麦克风为中心的圆弧上,与麦克风的距离固定在1至10米区间内,且在事件持续时长内以恒定角速度运动。由于采用了Ambisonic空间录音格式,远场平面波声源或点声源的导向矢量与频率无关,因此当声源移动时,无需使用时变卷积或冲激响应插值方案:单声道信号的空间编码是通过针对移动声源当前到达方向(Direction of Arrival, DOA)的瞬时Ambisonic编码矢量逐样本完成的。 数据集中的合成轨迹在方位角与俯仰角维度均存在变化,模拟得到的恒定角速度范围为[-90, 90]度/秒,步长为10度/秒。 本数据集的授权协议可在LICENSE文件中查阅。其余9个压缩包均为对应重叠数量与交叉验证划分的数据集。例如,ov3_split1.zip包含最多3个时域重叠声事件(ov3)与第1组交叉验证划分(split1)场景下的音频文件夹与元数据文件夹。在每个音频/元数据文件夹中,训练划分的文件名带有"train"前缀,测试划分的文件名则带有"test"前缀。 本数据集是作为《基于卷积循环神经网络的多移动声源定位、检测与跟踪》研究工作的一部分收集完成的。




