遇见数据集

Test dataset for separation of speech, traffic sounds, wind noise, and general sounds

收藏
Zenodo2021-03-24 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

The dataset was generated as part of the paper:<br> Deep Complex U-Net Ensemble for Outdoor Urban Sound Source Separation,<br> K. Arendt, A. Szumaczuk, B. Jasik, P. Masztalski, K. Piaskowski, M. Matuszewski, K. Nowicki, P. Zborowski. It contains various sounds from the Audio Set [1] and spoken utterances from VCTK [2] and DNS [3] datasets. Contents:<br> sr_8k/<br> mix_clean/<br> s1/<br> s2/<br> s3/<br> s4/<br> sr_16k/<br> mix_clean/<br> s1/<br> s2/<br> s3/<br> s4/<br> sr_48k/<br> mix_clean/<br> s1/<br> s2/<br> s3/<br> s4/ Each directory contains 512 audio samples in different sampling rate (sr_8k - 8 kHz, sr_16k - 16 kHz, sr_48k - 48 kHz).<br> The audio samples for each sampling rate are different as they were generated randomly and separately.<br> Each directory contains 5 subdirectories:<br> - mix_clean - mixed sources,<br> - s1 - source #1 (general sounds),<br> - s2 - source #2 (speech),<br> - s3 - source #3 (traffic sounds),<br> - s4 - source #4 (wind noise). The sound mixtures were generated by adding s2, s3, s4 to s1 with SNR ranging from -10 to 10 dB w.r.t. s1. <br> REFERENCES: [1] Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman,<br> Aren Jansen, Wade Lawrence, R. Channing Moore,<br> Manoj Plakal, and Marvin Ritter, “Audio set: An ontology<br> and human-labeled dataset for audio events,” in<br> Proc. IEEE ICASSP 2017, New Orleans, LA, 2017. [2] Christophe Veaux, Junichi Yamagishi, and Kirsten Mac-<br> Donald, “CSTR VCTK corpus: English multi-speaker<br> corpus for CSTR voice cloning toolkit, [sound],”<br> https://doi.org/10.7488/ds/1994, University of Edinburgh.<br> The Centre for Speech Technology Research<br> (CSTR). 2017. [3] Chandan K. A. Reddy, Ebrahim Beyrami, Harishchandra<br> Dubey, Vishak Gopal, Roger Cheng, Ross Cutler,<br> Sergiy Matusevych, Robert Aichner, Ashkan Aazami,<br> Sebastian Braun, Puneet Rana, Sriram Srinivasan, and<br> Johannes Gehrke, “The interspeech 2020 deep noise<br> suppression challenge: Datasets, subjective speech<br> quality and testing framework,” 2020.

提供机构:
Zenodo
创建时间:
2020-11-18
二维码
社区交流群
二维码
科研交流群
商业服务