Artificial sound mixes with event insertions
收藏资源简介:
<strong>Contains artificial sound mixes and meta data that were created for the task of<br> sound event detection.</strong> The mixes were created using background and event audio recordings<br> from Tampere University's Detection and Classification of Acoustic Scenes and<br> Events (DCASE) Community. More information on the source data can be found at<br> http://www.cs.tut.fi/sgn/arg/dcase2017/challenge/task-rare-sound-event-detection#audio-dataset. Source data credits: Diment, Aleksandr et al (2017, 2018) <strong>Created artificial sound mixes are 10 seconds long and contain:</strong> - background audio from diverse scenes from start to end<br> - 0 to 4 event insertions With this formula, two datasets were created separately: a training, and an<br> evaluation dataset. These were created separately so that the backgrounds and<br> event recordings used for the evaluation dataset were not used in any of the<br> training audio mixes. Thus, keeping them unseen by the system during development. <strong>Audio data:</strong><br> <strong>mixes_train.zip</strong>: Contains the audio mixes created for training.<br> Audio format: .wav<br> Count: 1000 tracks <strong>mixes_eval.zip</strong>: Contains the audio mixes created for evaluation.<br> Audio format: .wav<br> Count: 500 tracks <strong>Meta-data:</strong><br> Meta-data was maintained documenting the source background, the overlaid source<br> events, and the time onset and offset of each sound event or confusing sound. <strong>meta_track_info_train.csv and meta_track_info_eval.csv columns:</strong> - trackID: the unique ID of a mix<br> - class_dummy: whether a mix contain a glas break event (other events can be<br> determined using meta_clip_insertions_train and meta_clip_insertions_eval)<br> - background_file: the unique file reference used for background sound to the<br> source data (i.e. DCASE original audio data set)<br> - background_t0: the second in the original background sound recording in which<br> the 10 second background starts. <strong>meta_clip_insertions_train.csv and meta_clip_insertions_eval.csv columns:</strong> - trackID: the unique ID of a mix<br> - event: the type of event insertion<br> - event_start: at which time in the mix the event start<br> - event_end: at which time in the mix the event ends<br> - event_file: the unique file_number of event reference to the<br> source data (i.e. DCASE original audio data set)<br> an event_file with value 345584_4.wav means the event comes from file<br> 345584.wav in the DCASE audio set and is the 4th event in that audio file. To see the project for which this data set was created visit the github<br> repository at https://github.com/reyvaz/sound-event-detection <strong>References:</strong><br> Diment, Aleksandr, Mesaros, Annamaria, Heittola, Toni, & Virtanen, Tuomas. (2017).<br> TUT Rare sound events, Development dataset [Data set]. Zenodo.<br> http://doi.org/10.5281/zenodo.401395 Aleksandr Diment, Annamaria Mesaros, Toni Heittola, & Tuomas Virtanen. (2018).<br> TUT Rare sound events, Evaluation dataset [Data set]. Zenodo.<br> http://doi.org/10.5281/zenodo.1160455<br>
**本数据集包含为声音事件检测任务构建的人工合成音频片段与元数据。** 该音频片段的合成素材来自坦佩雷大学声学场景与事件检测与分类(Detection and Classification of Acoustic Scenes and Events, DCASE)社区提供的背景音与事件音频录音。有关源数据的更多信息可访问:http://www.cs.tut.fi/sgn/arg/dcase2017/challenge/task-rare-sound-event-detection#audio-dataset。 源数据致谢:Diment, Aleksandr 等人(2017、2018) **所构建的人工合成音频片段均为10秒时长,包含以下内容:** - 贯穿全程的多场景背景音频 - 0至4段插入式事件音频 基于该合成规则,本数据集分别构建了训练集与评估集。二者的构建相互独立,即评估集所使用的背景音与事件录音,不会出现在任何训练音频片段中,以此确保模型在开发阶段无法接触到评估集数据。 **音频数据:** **mixes_train.zip**:包含为训练任务构建的音频合成片段。音频格式:.wav;总曲目数:1000条。 **mixes_eval.zip**:包含为评估任务构建的音频合成片段。音频格式:.wav;总曲目数:500条。 **元数据:** 元数据用于记录源背景音、叠加的事件音频,以及每一段声音事件或干扰音的起始与结束时间。 **meta_track_info_train.csv 与 meta_track_info_eval.csv 字段说明:** - trackID:音频片段的唯一标识符 - class_dummy:标记该音频片段是否包含玻璃破碎事件(其余事件类型可通过meta_clip_insertions_train与meta_clip_insertions_eval文件进行查询) - background_file:指向源数据中背景音的唯一文件引用(即DCASE原始音频数据集) - background_t0:原始背景音录音中,10秒背景片段的起始时刻(单位:秒) **meta_clip_insertions_train.csv 与 meta_clip_insertions_eval.csv 字段说明:** - trackID:音频片段的唯一标识符 - event:插入式事件的类型 - event_start:该事件在音频片段中的起始时刻(单位:秒) - event_end:该事件在音频片段中的结束时刻(单位:秒) - event_file:指向源数据中事件音频的唯一文件编号引用(即DCASE原始音频数据集)。例如,取值为345584_4.wav的event_file,表示该事件来自DCASE音频集中的345584.wav文件,且为该文件中的第4段事件音频。 如需了解本数据集所支撑的研究项目,可访问其GitHub仓库:https://github.com/reyvaz/sound-event-detection **参考文献:** Diment, Aleksandr、Mesaros, Annamaria、Heittola, Toni 与 Virtanen, Tuomas. (2017). TUT稀有声音事件:开发数据集 [数据集]. Zenodo. http://doi.org/10.5281/zenodo.401395 Aleksandr Diment、Annamaria Mesaros、Toni Heittola 与 Tuomas Virtanen. (2018). TUT稀有声音事件:评估数据集 [数据集]. Zenodo. http://doi.org/10.5281/zenodo.1160455



