SONYC-Backgrounds: a collection of urban background recordings from an acoustic sensor network
收藏资源简介:
<strong>Created by</strong> Aurora Cramer <sup>(1, 2)</sup>, Mark Cartwright <sup>(3)</sup>, Fatemeh Pishdadian <sup>(4)</sup>, Juan Pablo Bello <sup>(1,2,5,6)</sup> 1. Music and Audio Research Lab, New York University<br> 2. Department of Electrical and Computer Engineering, New York University<br> 3. Department of Informatics, New Jersey Institute of Technology<br> 4. Interactive Audio Lab, Northwestern University<br> 5. Center for Urban Science and Progress, New York University<br> 6. Department of Computer Science and Engineering, New York University <br> <strong>Publication</strong> If you use this data in your work, please cite the following paper, which introduced this dataset: [1] Cramer, A., Cartwright, M., Pishdadian, F., and Bello, J.P. Weakly Supervised Source-Specific Sound Level Estimation in Noisy Soundscapes. In Proceedings of the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2021. [pdf] <br> <strong>Description</strong> SONYC-Backgrounds is an open dataset of recordings of urban background noise obtained from the SONYC acoustic sensor network [2]. This dataset was developed with the goal of synthesizing soundscapes with a diverse set of realistic sounding background activity, for use in developing and evaluating machine listening systems in urban settings. <br> <strong>Data acquisition</strong> The provided audio has been acquired using the SONYC acoustic sensor network for urban noise pollution monitoring [2]. Over 50 different sensors have been deployed in New York City. All recordings are 10 seconds and were recorded with identical microphones at identical gain settings. <br> <strong>Recording selection</strong> From the large collection of audio recordings acquired in 2017, we obtain a much smaller subset of likely background recordings. We first process the dataset using a sensor fault detector to filter out recordings with artifacts caused by hardware failures in the sensors. The sensor fault detector is a random forest, trained with a small collection of audio examples using active learning [3]. We then determine if a recording is background or not using an urban sound classifier trained to detect the presence of sources of interest to urban noise pollution monitoring [4, 5]. We use the classifier to find recordings that *do not* contain the sound classes of interest. The classifier model is a multi-layer perception with two hidden layers, which takes as input an OpenL3 embedding [6] for a 1 s clip of audio and produces multi-label prediction probabilities for each class. This model is nearly identical to the one used for the DCASE 2019 Challenge Urban Sound Tagging Task baseline model, aside from the addition of an extra hidden layer. Predictions for entire recordings are obtained by max-pooling the predictions for each class across time. A recording is considered background if the probabilities of the target classes fall below their respective detection thresholds, i.e. no target classes are detected. The classifier was trained on the SONYC-UST v1 dataset [4], and the detection thresholds for each class were tuned to correspond to 70% <em>negative</em> recall (true negative rate) on the test set to increase the likelihood that recordings are background. After this selection process, we obtain 441 background clips. <strong>Metadata</strong> To maintain privacy, the recordings in this release have been distributed in time and location, and recording times have been quantized to the hour. Sensor IDs are consistent with those SONYC-UST dataset [4]. The corresponding location of the sensors can be found in the SONYC-UST v2 dataset [5], though these locations have been mapped to the "block" level to maintain privacy. See the DCASE 2020 Challenge Urban Sound Tagging with Spatiotemporal Context Task page for more information on the metadata. <br> <strong>Data splits</strong> The dataset is partitioned into a train/valid/test split of roughly 60/20/20, using a simple greedy method to assign sensors to subsets. <br> <strong>Files</strong> The dataset directory contains the directories `train`, `valid`, and `test` for each of the respective data subsets. Each directory contains recordings, with the file format: `<sensor-id>_<year>-<month>-<day>_<hour>_<instance-num>.wav`, where `<instance-num>` is used to distinguish recordings from the same sensor occurring during the same hour. Aside from `<year>`, each of these fields in the format are lead zero padded to two places (i.e. `printf` format `"%02d"`). <br> <strong>Conditions of use</strong> Dataset created by Aurora Cramer, Mark Cartwright, Fatemeh Pishdadian, and Juan Pablo Bello. The SONYC-Backgrounds dataset is offered free of charge under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) license: https://creativecommons.org/licenses/by/4.0/ The dataset and its contents are made available on an “as is” basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, New York University is not liable for, and expressly excludes all liability for, loss or damage however and whenever caused to anyone by any use of the SONYC-Backgrounds dataset or any part of it. <strong>Contact</strong> If you have any questions, comments, or concerns, please direct correspondence to Aurora Cramer (aurora (dot) linh (dot) cramer (at) gmail (dot) com). <br> <strong>References and Links</strong> [1] Cramer, A., Cartwright, M., Pishdadian, F., and Bello, J.P. Weakly Supervised Source-Specific Sound Level Estimation in Noisy Soundscapes. In Proceedings of the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2021. [2] Bello, J. P., Silva, C., Nov, O., Dubois, R. L., Arora, A., Salamon, J., C. Mydlarz, and Doraiswamy, H. (2019). Sonyc: A system for monitoring, analyzing, and mitigating urban noise pollution. Communications of the ACM, 62(2), 68-77. [3] Wang, Y., Mendez, A.E.M., Cartwright, M., and Bello, J.P. Active Learning for Efficient Audio Annotation and Classification with a Large Amount of Unlabeled Data. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019. [4] Cartwright, M., Mendez, A.E.M., Cramer, A., Lostanlen, V., Dove, G., Wu, H., Salamon, J., Nov, O., and Bello, J.P. SONYC Urban Sound Tagging (SONYC-UST): A Multilabel Dataset from an Urban Acoustic Sensor Network. In Proceedings of the Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE) , 2019. [5] Cartwright, M., Cramer, A., Mendez, A.E.M., Wang, Y., Wu, H., Lostanlen, V., Fuentes, M., Dove, G., Mydlarz, C., Salamon, J., Nov, O., and Bello, J.P. SONYC-UST-V2: An Urban Sound Tagging Dataset with Spatiotemporal Context. In Proceedings of the Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE), 2020. [6] Look, Listen and Learn More: Design Choices for Deep Audio Embeddings<br> Cramer, A., Wu, H.-H., Salamon J., and Bello. J.P. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019. <br> <strong>Acknowledgements</strong> We would like to thank all those involved in the SONYC project. This work is partially supported by National Science Foundation award 1633259 and award 1544753.
### 数据集创建者 本数据集由Aurora Cramer <sup>(1, 2)</sup>、Mark Cartwright <sup>(3)</sup>、Fatemeh Pishdadian <sup>(4)</sup>与Juan Pablo Bello <sup>(1,2,5,6)</sup> 创建。 1. 纽约大学音乐与音频研究实验室(Music and Audio Research Lab, New York University) 2. 纽约大学电气与计算机工程系(Department of Electrical and Computer Engineering, New York University) 3. 新泽西理工学院信息学系(Department of Informatics, New Jersey Institute of Technology) 4. 西北大学交互音频实验室(Interactive Audio Lab, Northwestern University) 5. 纽约大学城市科学与进步中心(Center for Urban Science and Progress, New York University) 6. 纽约大学计算机科学与工程系(Department of Computer Science and Engineering, New York University) ### 引用规范 若您在研究工作中使用本数据集,请引用下述介绍该数据集的论文: [1] Cramer, A., Cartwright, M., Pishdadian, F., 与 Bello, J.P. 噪声声景中的弱监督源特定声级估计. 发表于IEEE信号处理与音频声学应用研讨会(IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, WASPAA), 2021. [PDF] ### 数据集概述 SONYC-Backgrounds是从SONYC声学传感器网络[2]获取的城市背景噪声录音开源数据集。本数据集旨在合成包含多样化真实背景活动的声景,用于开发与评估城市环境下的机器听觉系统。 ### 数据采集 本数据集所提供的音频通过用于城市噪声污染监测的SONYC声学传感器网络[2]采集。纽约市已部署超过50个此类传感器,所有录音时长均为10秒,且采用统一型号的麦克风与相同增益设置录制。 ### 录音筛选流程 从2017年采集的海量音频录音中,我们筛选出一小部分符合要求的背景录音。首先,我们使用传感器故障检测器处理数据集,过滤掉因传感器硬件故障产生伪影的录音。该传感器故障检测器为随机森林(random forest),通过主动学习(active learning)[3]使用少量音频样本训练得到。随后,我们使用经训练的城市声音分类器判断录音是否为背景音,该分类器用于检测城市噪声污染监测关注的声源[4,5]。我们利用该分类器筛选出**不**包含目标声音类别的录音。该分类器模型为包含两个隐藏层的多层感知器(multi-layer perception),输入为1秒音频片段的OpenL3嵌入(OpenL3 embedding)[6],并为每个类别输出多标签预测概率。除额外增加一个隐藏层外,该模型与DCASE 2019挑战赛城市声音标记任务的基线模型基本一致。通过对每类的时域预测结果进行最大池化(max-pooling),可得到整条录音的预测结果。若目标类别的概率低于各自的检测阈值,即未检测到任何目标类别,则该录音被认定为背景音。本分类器基于SONYC-UST v1数据集[4]训练得到,且将各类别的检测阈值调整为在测试集上达到70%的**负召回率**(negative recall,即真阴性率),以提高录音为背景音的可能性。经过上述筛选流程后,最终得到441条背景音频片段。 ### 元数据说明 为保护隐私,本次发布的录音在时间与空间分布上已做脱敏处理,录音时间已按小时量化。传感器ID与SONYC-UST数据集[4]保持一致。传感器的对应位置可在SONYC-UST v2数据集[5]中查询,但为保护隐私,这些位置已映射至“街区”级别。有关元数据的更多信息,请参见DCASE 2020挑战赛基于时空上下文的城市声音标记任务页面。 ### 数据集划分 本数据集采用简单贪心方法将传感器分配至各子集,划分为训练集/验证集/测试集,比例约为60/20/20。 ### 文件结构 数据集目录包含分别对应各数据子集的`train`(训练集)、`valid`(验证集)与`test`(测试集)文件夹。每个文件夹内包含录音文件,文件名格式为:`<sensor-id>_<year>-<month>-<day>_<hour>_<instance-num>.wav`,其中`<instance-num>`用于区分同一传感器在同一小时内录制的多条录音。除年份外,格式中的其余字段均采用两位前导零填充,即遵循C语言`printf`格式"%02d"的填充规则。 ### 使用许可条款 本数据集由Aurora Cramer、Mark Cartwright、Fatemeh Pishdadian与Juan Pablo Bello创建。SONYC-Backgrounds数据集依据知识共享署名4.0国际许可协议(Creative Commons Attribution 4.0 International, CC BY 4.0)免费发布,协议链接:https://creativecommons.org/licenses/by/4.0/。本数据集及内容按“现状”提供,不附带任何形式的明示或默示担保,包括但不限于对合格质量、适销性、特定用途适用性、准确性或完整性,以及无错误的默示担保。除法律不可排除或限制的责任外,纽约大学明确排除因任何使用本数据集或其部分内容而对任何人造成的任何损失或损害的全部责任。 ### 联系方式 若您有任何疑问、评论或顾虑,请联系Aurora Cramer(邮箱:aurora.linh.cramer@gmail.com,原文标注为`aurora (dot) linh (dot) cramer (at) gmail (dot) com`)。 ### 参考文献与链接 [1] Cramer, A., Cartwright, M., Pishdadian, F., 与 Bello, J.P. Weakly Supervised Source-Specific Sound Level Estimation in Noisy Soundscapes. 发表于IEEE信号处理与音频声学应用研讨会(IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, WASPAA), 2021. [2] Bello, J. P., Silva, C., Nov, O., Dubois, R. L., Arora, A., Salamon, J., Mydlarz, C., 与 Doraiswamy, H. Sonyc: A system for monitoring, analyzing, and mitigating urban noise pollution. *Communications of the ACM*, 62(2), 68-77. 2019. [3] Wang, Y., Mendez, A.E.M., Cartwright, M., 与 Bello, J.P. 面向大量未标记数据的高效音频标注与分类的主动学习. 发表于IEEE国际声学、语音与信号处理会议(IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP), 2019. [4] Cartwright, M., Mendez, A.E.M., Cramer, A., Lostanlen, V., Dove, G., Wu, H., Salamon, J., Nov, O., 与 Bello, J.P. SONYC Urban Sound Tagging (SONYC-UST): 来自城市声学传感器网络的多标签数据集. 发表于检测与分类声学场景与事件研讨会(Workshop on Detection and Classification of Acoustic Scenes and Events, DCASE), 2019. [5] Cartwright, M., Cramer, A., Mendez, A.E.M., Wang, Y., Wu, H., Lostanlen, V., Fuentes, M., Dove, G., Mydlarz, C., Salamon, J., Nov, O., 与 Bello, J.P. SONYC-UST-V2: 包含时空上下文的城市声音标记数据集. 发表于检测与分类声学场景与事件研讨会(Workshop on Detection and Classification of Acoustic Scenes and Events, DCASE), 2020. [6] Cramer, A., Wu, H.-H., Salamon J., 与 Bello, J.P. 观其形,闻其声,学得更多:深度音频嵌入的设计选择. 发表于IEEE国际声学、语音与信号处理会议(IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP), 2019. ### 致谢 感谢SONYC项目全体参与人员。本研究部分得到美国国家科学基金会(National Science Foundation)编号1633259与1544753的项目资助。



