[DCASE2024 Task 3] Synthetic SELD mixtures for baseline training
收藏资源简介:
DESCRIPTION:This audio dataset serves serves as supplementary material for the DCASE2024 Challenge Task 3: Audio and Audiovisual Sound Event Localization and Detection with Distance Estimation. The dataset consists of synthetic spatial audio mixtures of sound events spatialized for two different spatial formats using real measured room impulse responses (RIRs) measured in various spaces of Tampere University (TAU). The mixtures are generated using the same process as the one used to generate the recordings of the TAU-NIGENS Spatial Sound Scenes 2021 dataset for the DCASE2021 Challenge Task 3. The SELD task setup in DCASE2024 is based on spatial recordings of real scenes, captured in the STARS23 dataset. Since the task setup allows use of external data, these synthetic mixtures serve as additional training material for the baseline model. For more details on the task setup, please refer to the task description. Note that the generator code and the collection of room responses used to spatialize sound samples will be also be made available soon. For more details on the recording of RIRs, spatialization, and generation, see: Archontis Politis, Sharath Adavanne, Daniel Krause, Antoine Deleforge, Prerak Srivastava, Tuomas Virtanen (2021). A Dataset of Dynamic Reverberant Sound Scenes with Directional Interferers for Sound Event Localization and Detection. In Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2021), Barcelona, Spain. available here. SPECIFICATIONS: 13 target sound classes (see task description for details) The sound event samples are sources from the FSD50K dataset, based on affinity of the labels in that dataset to the target classes. The selection on distinguishing which labels in FSD50K corresponded to the target ones, then selecting samples that were tagged with only those labels, and additionally that they had annotator rating of Present and Predominant (see FSD50K for more details). The list of the selected files is included here. 1200 1-minute long spatial recordings Sampling rate of 24kHz Two 4-channel recording formats, first-order Ambisonics (FOA) and tetrahedral microphone array (MIC) Spatial events spatialized in 9 unique rooms, using measured RIRs for the two formats Maximum polyphony of 3 (with possible same-class events overlapping) Even though the whole set is used for training of the baseline without distinction between the mixtures, we have included a separation into a training and testing split, in case on one needs to test the performance purely on those synthetic conditions (for example for comparisons with training on mixed synthetic-real data, fine-tuning on real data, or training on real data only). The training split is indicated as fold1 in the dataset, contains 900 recordings spatialized on 6 rooms (150 recordings/room) and it is based on samples from the development set of FSD50K. The testing split is indicated as fold2 in the dataset, contains 300 recordings spatialized on 3 rooms (100 recordings/room) and it is based on samples from the evaluation set of FSD50K. Common metadata files for both formats are provided. For the file naming and the metadata format, refer to the task setup. DOWNLOAD INSTRUCTIONS: Download the zip files and use your preferred compression tool to unzip these split zip files. To extract a split zip archive (named as zip, z01, z02, ...), you could use, for example, the following syntax in Linux or OSX terminal: Combine the split archive to a single archive: zip -s 0 split.zip --out single.zip Extract the single archive using unzip: unzip single.zip
数据集说明: 本音频数据集为DCASE2024挑战赛任务3(带距离估计的音频与视听声事件定位与检测)的补充材料。本数据集包含基于真实测量的房间冲激响应(Room Impulse Responses, RIRs)生成的合成空间音频混合片段,这些声事件采用两种不同空间格式进行空间化处理,房间冲激响应采集自坦佩雷大学(Tampere University, TAU)的多个场地。该混合片段的生成流程与DCASE2021挑战赛任务3所用的TAU-NIGENS 2021空间声场景数据集的录制生成流程完全一致。 DCASE2024的声事件定位与检测(SELD)任务设置基于STARS23数据集采集的真实场景空间录音。由于本次任务允许使用外部数据,此类合成混合片段可作为基线模型的额外训练素材。如需了解任务设置的更多细节,请参阅任务说明文档。 请注意,用于声事件样本空间化处理的生成代码与房间冲激响应集即将公开。如需了解房间冲激响应采集、空间化处理及数据生成的更多细节,请参阅以下文献: Archontis Politis、Sharath Adavanne、Daniel Krause、Antoine Deleforge、Prerak Srivastava、Tuomas Virtanen(2021)。面向声事件定位与检测的带定向干扰源的动态混响声场景数据集。收录于2020年声学场景检测与分类研讨会(DCASE2021)论文集,西班牙巴塞罗那。 公开链接见此处。 数据集规格: 共包含13个目标声类别(详细分类规则请参见任务说明文档)。 本数据集的声事件样本源自FSD50K数据集,筛选依据为该数据集的标签与本次任务目标类别的匹配度。具体筛选流程为:首先识别FSD50K中与目标类别匹配的标签,随后仅保留被标注为这些标签的样本,且额外要求样本的标注者评分需为“Present(存在)”与“Predominant(主要)”(详细说明请参阅FSD50K数据集文档)。所选样本的文件清单已附于本数据集内。 共1200条时长为1分钟的空间录音。 采样率为24kHz。 支持两种4通道录音格式:一阶Ambisonics(FOA)与四面体麦克风阵列(MIC)。 空间事件的空间化处理基于9个独立房间的实测房间冲激响应,且两种录音格式均使用对应采集的冲激响应。 最大复音数为3(允许同类声事件重叠出现)。 尽管基线模型的训练可直接使用全部混合片段而无需区分,但本数据集仍划分了训练集与测试集,以供研究者仅基于合成数据场景测试模型性能(例如用于对比合成-真实混合数据训练、真实数据微调或仅基于真实数据训练的实验效果)。 训练集在数据集中以fold1标识,包含900条录音,分别来自6个房间(每个房间150条录音),样本源自FSD50K数据集的开发集。 测试集在数据集中以fold2标识,包含300条录音,分别来自3个房间(每个房间100条录音),样本源自FSD50K数据集的评估集。 两种录音格式共享通用元数据文件。关于文件命名规则与元数据格式,请参阅任务设置文档。 下载说明: 下载分卷压缩包并使用任意解压工具进行解压。若需在Linux或OSX终端中合并并解压分卷归档(文件命名格式为zip、z01、z02……),可使用如下命令: 将分卷归档合并为单文件归档: zip -s 0 split.zip --out single.zip 使用unzip命令解压单文件归档: unzip single.zip



