Seeing Sound Dataset
收藏资源简介:
该数据集包含合成的声景和众包音频事件注释,用于研究声景复杂性和声音可视化对音频事件注释质量与速度的影响。数据集通过两个维度(最大复音和Gini复音)变化声景复杂度,共合成60个10秒长的声景,每个声景由90名Amazon Mechanical Turk参与者注释,其中30名使用波形可视化,30名使用频谱图可视化,30名无任何可视化辅助。
This dataset comprises synthesized soundscapes and crowdsourced audio event annotations, designed to investigate the impact of soundscape complexity and sound visualization on the quality and speed of audio event annotation. The dataset varies soundscape complexity along two dimensions (maximum polyphony and Gini polyphony), synthesizing a total of 60 soundscapes, each 10 seconds in length. Each soundscape was annotated by 90 Amazon Mechanical Turk participants, with 30 using waveform visualization, 30 using spectrogram visualization, and 30 without any visualization aids.
数据集概述
创建者
- 机构: 纽约大学(美国)、滑铁卢大学(加拿大)
- 作者: Mark Cartwright, Ayanna Seals, Justin Salamon, Alex Williams, Stefanie Mikloska, Duncan MacConnell, Edith Law, Juan Pablo Bello, Oded Nov
描述
- 目的: 研究声音景观复杂性和声音可视化对声音事件(如开始时间、结束时间、声音类别和接近度)注释质量和速度的影响。
- 方法: 通过两个维度(最大复调性和吉尼复调性)变化声音景观复杂性,合成60个10秒长的声音景观,并由90名Amazon Mechanical Turk参与者进行注释。
内容
- 音频文件: 60个WAV格式的声音景观文件,命名格式为
soundscape-<soundscape_id>_m<max_polyphony_level>_g<gini_polyphony_level>.wav。 - 注释文件: 60个JAMS格式的注释文件,命名格式为
soundscape-<soundscape_id>_m<max_polyphony_level>_g<gini_polyphony_level>.jams。
注释文件详情
- 格式: JAMS,一种JSON基础的音频注释格式。
- 内容: 每个JAMS文件包含地面实况注释和90个众包注释。
- 地面实况注释: 描述声音事件的时间、持续时间和值,包括声音类别、信噪比等。
- 众包注释: 描述参与者感知的开始时间、持续时间和值,包括感知的声音类别和接近度。
声音事件类别
- 汽车喇叭声
- 狗叫声
- 引擎怠速声
- 枪声
- 钻机声
- 音乐播放声
- 人群呼喊声
- 人群交谈声
- 警笛声
引用信息
-
BibTeX引用:
@article{Cartwright:SeeingSound:CSCW:17, Author = {Cartwright, M. and Seals, A. and Salamon, J. and Williams, A. and Mikloska, S. and MacConnell, D. and Law, E. and Bello, J.P. and Nov, O.}, Journal = {Proceedings of the ACM on Human-Computer Interaction}, Number = {2}, Title = {Seeing Sound: Investigating the Effects of Visualizations and Complexity on Crowdsourced Audio Annotations}, Volume = {1}, Year = {2017}, DOI = {10.1145/3134664} }




