MAESTRO Real - Multi-Annotator Estimated Strong Labels
收藏资源简介:
The dataset was created for studying estimation of strong labels using crowdsourcing. It contains 49 real-life audio files from 5 different acoustic scenes, and the annotation outcome. Annotation was performed using Amazon Mechanical Turk. Total duration of the dataset is 189 minutes and 52 seconds Audio files are a subset from TUT Acoustic Scenes 2016 dataset, belonging to five acoustic scenes: cafe/restaurant, city center, grocery store, metro station and residential area. Each scene have 6 classes, some of them are common to all the scenes, resulting into 17 classes in total. <br> The dataset contains: audio: the 49 real-life recordings, each from 3 to 5 min long. soft labels: estimated strong labels from the crowdsourced data, values between 0 and 1 indicates the uncertainty of the annotators. For more details about the real-life recordings, please see the following paper: A. Mesaros, T. Heittola and T. Virtanen, "TUT database for acoustic scene classification and sound event detection," <em>2016 24th European Signal Processing Conference (EUSIPCO)</em>, 2016, pp. 1128-1132.
本数据集专为基于众包的强标签(strong labels)估计研究而构建,包含来自5种不同声学场景的49条真实音频文件与标注结果。本次标注通过Amazon Mechanical Turk完成。数据集总时长为189分52秒,所用音频文件均为TUT Acoustic Scenes 2016数据集的子集,涵盖5类声学场景:咖啡馆/餐厅、城市中心、杂货店、地铁站与居民区。每个场景下设6个类别,其中部分类别为所有场景共有,最终总计17个类别。 本数据集包含以下内容: 1. 音频:49条真实录音,单条时长为3至5分钟。 2. 软标签(soft labels):基于众包数据估算得到的强标签(strong labels),取值介于0至1之间,用于反映标注人员的标注不确定性。 若需了解更多关于该真实录音的细节,请参阅以下论文:A. Mesaros、T. Heittola与T. Virtanen, "TUT database for acoustic scene classification and sound event detection", *2016年第24届欧洲信号处理大会(EUSIPCO)*, 2016, pp. 1128-1132.



