遇见数据集

Urban Sound & Sight (Urbansas) - Labeled set

收藏
Zenodo2022-06-19 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>Urban Sound &amp; Sight (Urbansas): </strong> Version 1.0, May 2022 <strong>Created by</strong><br> Magdalena Fuentes (1, 2), Bea Steers (1, 2), Pablo Zinemanas (3), Martín Rocamora (4), Luca Bondi (5), Julia Wilkins (1, 2), Qianyi Shi (2), Yao Hou (2), Samarjit Das (5), Xavier Serra (3), Juan Pablo Bello (1, 2)<br> 1. Music and Audio Research Lab, New York University<br> 2. Center for Urban Science and Progress, New York University<br> 3. Universitat Pompeu Fabra, Barcelona, Spain<br> 4. Universidad de la República, Montevideo, Uruguay<br> 5. Bosch Research, Pittsburgh, PA, USA <strong>Publication</strong> If using this data in academic work, please cite the following paper, which presented this dataset:<br> M. Fuentes, B. Steers, P. Zinemanas, M. Rocamora, L. Bondi, J. Wilkins, Q. Shi, Y. Hou, S. Das, X. Serra, J. Bello. “Urban Sound &amp; Sight: Dataset and Benchmark for Audio-Visual Urban Scene Understanding”. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022. <strong>Description</strong> Urbansas is a dataset for the development and evaluation of machine listening systems for audiovisual spatial urban understanding. One of the main challenges to this field of study is a lack of realistic, labeled data to train and evaluate models on their ability to localize using a combination of audio and video.<br> We set four main goals for creating this dataset: <br> 1. To compile a set of real-field audio-visual recordings;<br> 2. The recordings should be stereo to allow exploring sound localization in the wild;<br> 3. The compilation should be varied in terms of scenes and recording conditions to be meaningful for training and evaluation of machine learning models;<br> 4. The labeled collection should be accompanied by a bigger unlabeled collection with similar characteristics to allow exploring self-supervised learning in urban contexts.<br> Audiovisual data<br> We have compiled and manually annotated Urbansas from two publicly available datasets, plus the addition of unreleased material. The public datasets are the TAU Urban Audio-Visual Scenes 2021 Development dataset (street-traffic subset) and the Montevideo Audio-Visual Dataset (MAVD): <br> Wang, Shanshan, et al. "A curated dataset of urban scenes for audio-visual scene analysis." ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021. Zinemanas, Pablo, Pablo Cancela, and Martín Rocamora. "MAVD: A dataset for sound event detection in urban environments." Detection and Classification of Acoustic Scenes and Events, DCASE 2019, New York, NY, USA, 25–26 oct, page 263--267 (2019). <br> The TAU dataset consists of 10-second segments of audio and video from different scenes across European cities, traffic being one of the scenes. Only the scenes labeled as traffic were included in Urbansas. MAVD is an audio-visual traffic dataset curated in different locations of Montevideo, Uruguay, with annotations of vehicles and vehicle components sounds (e.g. engine, brakes) for sound event detection. Besides the published datasets, we include a total of 9.5 hours of unpublished material recorded in Montevideo, with the same recording devices of MAVD but including new locations and scenes. Recordings for TAU were acquired using a GoPro Hero 5 (30fps, 1280x720) and a Soundman OKM II Klassik/studio A3 electret binaural in-ear microphone with a Zoom F8 audio recorder (48kHz, 24 bits, stereo). Recordings for MAVD were collected using a GoPro Hero 3 (24fps, 1920x1080) and a SONY PCM-D50 recorder (48kHz, 24 bits, stereo). When compiled in Urbansas, it includes 15 hours of stereo audio and video, stored in separate 10 second MPEG4 (1280x720, 24fps) and WAV (48kHz, 24 bit, 2 channel) files. Both released video datasets are already anonymized to obscure people and license plates, the unpublished MAVD data was anonymized similarly using this anonymizer. We also distribute the 2fps video used for producing the annotations. The audio and video files both share the same filename stem, meaning that they can be associated after removing the parent directory and extension. MAVD:<br> video/&lt;location_id&gt;_&lt;mavd_clip_id&gt;_&lt;clip_split_id&gt;.mp4<br> audio/&lt;location_id&gt;_&lt;mavd_clip_id&gt;_&lt;clip_split_id&gt;.wav TAU:<br> video/&lt;location_id&gt;_&lt;tau_clip_id&gt;.mp4<br> audio/&lt;location_id&gt;_&lt;tau_clip_id&gt;.wav <br> where location_id in both cases includes the city and an ID number. <br> city &amp; places &amp; clips &amp; mins &amp; frames &amp; labeled mins \\<br> Montevideo &amp; 8 &amp; 4085 &amp; 681 &amp; 980400 &amp; 92 \\<br> Stockholm &amp; 3 &amp; 91 &amp; 15 &amp; 21840 &amp; 2 \\<br> Barcelona &amp; 4 &amp; 144 &amp; 24 &amp; 34560 &amp; 24 \\<br> Helsinki &amp; 4 &amp; 144 &amp; 24 &amp; 34560 &amp; 16 \\<br> Lisbon &amp; 4 &amp; 144 &amp; 24 &amp; 34560 &amp; 19 \\<br> Lyon &amp; 4 &amp; 144 &amp; 24 &amp; 34560 &amp; 6 \\<br> Paris &amp; 4 &amp; 144 &amp; 24 &amp; 34560 &amp; 2 \\<br> Prague &amp; 4 &amp; 144 &amp; 24 &amp; 34560 &amp; 2 \\<br> Vienna &amp; 4 &amp; 144 &amp; 24 &amp; 34560 &amp; 6 \\<br> London &amp; 5 &amp; 144 &amp; 24 &amp; 34560 &amp; 4 \\<br> Milan &amp; 6 &amp; 144 &amp; 24 &amp; 34560 &amp; 6 \\<br> \midrule<br> Total &amp; 50 &amp; 5472 &amp; 912 &amp; 1.3M &amp; 180 \\ <br> <strong>Annotations</strong> <br> Of the 15 hours of audio and video, 3 hours of data (1.5 hours TAU, 1.5 hours MAVD) are manually annotated by our team both in audio and image, along with 12 hours of unlabeled data (2.5 hours TAU, 9.5 hours of unpublished material) for the benefit of unsupervised models. The distribution of clips across locations was selected to maximize variance across different scenes. The annotations were collected at 2 frames per second (FPS) as it provided a balance between temporal granularity and clip coverage. The annotation data is contained in video_annotations.csv and audio_annotations.csv. <strong>Video Annotations</strong> Each row in the video annotations represents a single object in a single frame of the video. The annotation schema is as follows: frame_id: The index of the frame within the clip the annotation is associated with. This index is 0-based and goes up to 19 (assuming 10-second clips with annotations at 2 FPS) track_id: The ID of the detected instance that identifies the same object across different frames. These IDs are guaranteed to be unique within a clip. x, y, w, h: The top-left corner and width and height of the object’s bounding box in the video. The values are given in absolute coordinates with respect to the image size (1280x720). class_id: The index of the class corresponding to: [0, 1, 2, 3, -1] — see label for the index mapping. The -1 value corresponds to the case where there are no events, but still clip-level annotations, like night and city. When operating on bounding boxes, class_id of -1 should be filtered. label: The label text. This is equivalent to LABELS[class_id], where LABELS=[car, bus, motorbike, truck, -1]. The label -1 has the same role as above. visibility: The visibility of the object. This is 1 unless the object becomes obstructed, where it changes to 0. filename: The file ID of the associated file. This is the file’s path minus the parent directory and extension. city: The city where the clip was collected in. location_id: The specific name of the location. This may include an integer ID following the city name for cases where there are multiple collection points. time: The time (in seconds) of the annotation, relative to the start of the file. Equivalent to frame_id / fps . night: Whether the clip takes place during the day or at night. This value is singular per clip. subset: Which data source the data originally belongs to (TAU or MAVD). <strong>Audio Annotations</strong> Each row represents a single object instance, along with the time range that it exists within the clip. The annotation schema is as follows: filename: The file ID odd the associated audio file. See filename above. class_id, label: See above. Audio has an additional class_id of 4 (label=offscreen) which indicates an off-screen vehicle - meaning a vehicle that is heard but not seen. A class_id of -1 indicates a clip-level annotation for a clip that has no object annotations (an empty scene). non_identifiable_vehicle_sound: True if the region contains the sound of vehicles where individual instances cannot be uniquely identified. start, end: The start and end times (in seconds) of the annotation relative to the file. <strong>Conditions of use</strong> Dataset created by Magdalena Fuentes, Bea Steers, Pablo Zinemanas, Martín Rocamora, Luca Bondi, Julia Wilkins, Qianyi Shi, Yao Hou, Samarjit Das, Xavier Serra, and Juan Pablo Bello. The Urbansas dataset is offered free of charge under the following terms: Urbansas annotations are release under the CC BY 4.0 license Urbansas video and audio replicates the original sources licenses: MAVD subset is released under CC BY 4.0 TAU subset is released under a Non-Commercial license <strong>Feedback</strong> Please help us improve Urbansas by sending your feedback to: Magdalena Fuentes: mfuentes@nyu.edu Bea Steers: bsteers@nyu.edu In case of a problem, please include as many details as possible. <strong>Acknowledgments</strong> This work was partially supported by the National Science Foundation award 1955357 and Bosch RTC.

**城市声音与视觉(Urbansas):** 版本1.0,2022年5月 **创建者** Magdalena Fuentes(1, 2)、Bea Steers(1, 2)、Pablo Zinemanas(3)、Martín Rocamora(4)、Luca Bondi(5)、Julia Wilkins(1, 2)、Qianyi Shi(2)、Yao Hou(2)、Samarjit Das(5)、Xavier Serra(3)、Juan Pablo Bello(1, 2) 1. 纽约大学音乐与音频研究实验室 2. 纽约大学城市科学与进步中心 3. 西班牙巴塞罗那庞培法布拉大学 4. 乌拉圭蒙得维的亚共和国大学 5. 美国宾夕法尼亚州匹兹堡博世研究院 **引用要求** 若将本数据集用于学术研究,请引用如下发表本数据集的论文: M. Fuentes, B. Steers, P. Zinemanas, M. Rocamora, L. Bondi, J. Wilkins, Q. Shi, Y. Hou, S. Das, X. Serra, J. Bello. "城市声音与视觉:面向视听城市场景理解的数据集与基准测试集". 2022年IEEE国际声学、语音与信号处理会议(ICASSP). **数据集描述** Urbansas是一款用于开发与评估视听空间城市场景理解系统的机器听觉数据集。当前该研究领域的核心挑战之一,是缺乏贴合真实场景的标注数据,用以训练和评估模型结合音频与视觉信息进行声源定位的能力。 我们在构建本数据集时设定了四大核心目标: 1. 采集一系列真实场景下的音视频录制数据; 2. 采用立体声录制格式,以便开展野外环境下的声源定位研究; 3. 覆盖多样化的场景与录制环境,以满足机器学习模型训练与评估的实际需求; 4. 配套提供规模更大、特征相似的未标注数据集,以支持城市场景下的自监督学习研究。 **视听数据** 我们从两个公开数据集以及未公开的原始素材中整理并手动标注构建了Urbansas。公开数据集分别为TAU城市音视频场景2021开发数据集(道路交通场景子集)与蒙得维的亚音视频数据集(MAVD): - Wang, Shanshan, et al. "A curated dataset of urban scenes for audio-visual scene analysis." ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021. (译:王珊珊等,《面向音视频场景分析的城市场景精选数据集》,2021年IEEE国际声学、语音与信号处理会议(ICASSP),IEEE出版社,2021年) - Zinemanas, Pablo, Pablo Cancela, and Martín Rocamora. "MAVD: A dataset for sound event detection in urban environments." Detection and Classification of Acoustic Scenes and Events, DCASE 2019, New York, NY, USA, 25–26 oct, page 263--267 (2019). (译:Pablo Zinemanas、Pablo Cancela与Martín Rocamora,《MAVD:面向城市环境声音事件检测的数据集》,2019年声学场景与事件检测与分类研讨会(DCASE 2019),美国纽约,10月25-26日,第263-267页,2019年) TAU数据集包含来自欧洲多座城市不同场景的10秒音视频片段,道路交通为其中一类场景。Urbansas仅收录了其中标注为道路交通的场景。MAVD是在乌拉圭蒙得维的亚多个地点采集的道路交通音视频数据集,附带用于声音事件检测的车辆及其部件(如发动机、刹车)的声音标注。除上述公开数据集外,我们还补充了9.5小时未公开的蒙得维的亚本地录制素材,录制设备与MAVD一致,但覆盖了全新的采集地点与场景。 TAU数据集的录制采用GoPro Hero 5(帧率30fps,分辨率1280×720)与Soundman OKM II Klassik/studio A3驻极体双耳入耳式麦克风搭配Zoom F8音频录制设备(采样率48kHz,位深24比特,立体声)。MAVD数据集的录制采用GoPro Hero 3(帧率24fps,分辨率1920×1080)与SONY PCM-D50录制设备(采样率48kHz,位深24比特,立体声)。 整合后的Urbansas数据集共包含15小时立体声音视频数据,分别存储为独立的10秒MPEG4格式(1280×720,24fps)视频文件与WAV格式(48kHz,24比特,双声道)音频文件。两个公开视频数据集均已完成匿名化处理,以模糊人脸与车牌;未公开的MAVD衍生素材亦采用相同工具完成匿名化。我们同时提供用于生成标注的2fps版本视频。音视频文件共享相同的文件名主干,移除父目录与扩展名后即可完成关联。 MAVD数据集文件命名规则: - 视频文件:video/<location_id>_<mavd_clip_id>_<clip_split_id>.mp4 - 音频文件:audio/<location_id>_<mavd_clip_id>_<clip_split_id>.wav TAU数据集文件命名规则: - 视频文件:video/<location_id>_<tau_clip_id>.mp4 - 音频文件:audio/<location_id>_<tau_clip_id>.wav 其中,两种数据集的location_id均包含城市名称与编号信息。 | 城市 | 地点数 | 片段数 | 时长(分钟) | 总帧数 | 标注时长(分钟) | |--------------|--------|--------|--------------|----------|------------------| | 蒙得维的亚 | 8 | 4085 | 681 | 980400 | 92 | | 斯德哥尔摩 | 3 | 91 | 15 | 21840 | 2 | | 巴塞罗那 | 4 | 144 | 24 | 34560 | 24 | | 赫尔辛基 | 4 | 144 | 24 | 34560 | 16 | | 里斯本 | 4 | 144 | 24 | 34560 | 19 | | 里昂 | 4 | 144 | 24 | 34560 | 6 | | 巴黎 | 4 | 144 | 24 | 34560 | 2 | | 布拉格 | 4 | 144 | 24 | 34560 | 2 | | 维也纳 | 4 | 144 | 24 | 34560 | 6 | | 伦敦 | 5 | 144 | 24 | 34560 | 4 | | 米兰 | 6 | 144 | 24 | 34560 | 6 | | **总计** | 50 | 5472 | 912 | 130万 | 180 | **标注信息** 在全部15小时的音视频数据中,团队针对其中3小时的数据(1.5小时TAU数据集片段、1.5小时MAVD数据集片段)完成了音视频双重手动标注,同时保留了12小时的未标注数据(2.5小时TAU数据集片段、9.5小时未公开素材),以支持无监督模型研究。片段在各城市的分布经过优化,以最大化不同场景间的多样性。标注采用每秒2帧的采样率,在时间粒度与片段覆盖范围之间取得平衡。标注数据存储于video_annotations.csv与audio_annotations.csv文件中。 #### 视频标注 视频标注文件的每一行对应视频单帧中的单个目标对象,标注格式如下: - frame_id:标注关联的片段内帧索引,从0开始计数,最大为19(对应10秒片段、2fps的标注帧率) - track_id:检测实例的唯一ID,用于标识不同帧中的同一目标对象,在单个片段内保证全局唯一 - x, y, w, h:视频中目标边界框的左上角坐标与宽高,采用相对于图像分辨率(1280×720)的绝对坐标值 - class_id:类别索引,对应映射关系为[0, 1, 2, 3, -1],详见label字段的索引说明。其中class_id为-1的情况代表无事件发生,但仍存在片段级标注(如昼夜、城市场景),处理边界框时需过滤该类别 - label:标签文本,对应LABELS[class_id],其中LABELS数组为[car, bus, motorbike, truck, -1],即[小汽车、公共汽车、摩托车、卡车、无事件] - visibility:目标可见性,取值为1,仅当目标被遮挡时变为0 - filename:关联文件的ID,即移除父目录与扩展名后的文件路径 - city:片段采集所在城市 - location_id:具体采集地点名称,若同一城市存在多个采集点,则在城市名称后附加整数编号 - time:标注相对于文件起始的时间(秒),等价于frame_id / fps - night:标识片段采集时段为白天或夜间,每个片段仅存在一个取值 - subset:数据来源子集,分为TAU或MAVD #### 音频标注 音频标注文件的每一行对应单个目标实例,及其在片段内的存在时间范围,标注格式如下: - filename:关联音频文件的ID,详见上述filename字段说明 - class_id, label:同视频标注中的定义。音频标注额外包含class_id为4(标签为offscreen,即屏外车辆),代表仅可闻而不可见的车辆 - class_id为-1时,代表该片段无目标对象标注(空场景)的片段级标注 - non_identifiable_vehicle_sound:若该区域包含无法唯一识别单个实例的车辆声音,则取值为True - start, end:标注相对于文件起始的开始与结束时间(秒) **使用条款** 本数据集由Magdalena Fuentes、Bea Steers、Pablo Zinemanas、Martín Rocamora、Luca Bondi、Julia Wilkins、Qianyi Shi、Yao Hou、Samarjit Das、Xavier Serra与Juan Pablo Bello共同创建。Urbansas数据集免费提供,使用条款如下: 1. Urbansas标注数据采用CC BY 4.0协议发布 2. Urbansas音视频数据保留原始数据源的授权协议: - MAVD子集采用CC BY 4.0协议发布 - TAU子集采用非商业使用协议发布 **反馈与改进** 若您有任何改进建议或反馈,请发送至: - Magdalena Fuentes: mfuentes@nyu.edu - Bea Steers: bsteers@nyu.edu 若您遇到任何问题,请尽可能提供详细细节。 **致谢** 本研究部分受到美国国家科学基金会奖项1955357与博世研发中心(Bosch RTC)的支持。

提供机构:
Zenodo
创建时间:
2022-06-19
二维码
社区交流群
二维码
科研交流群
商业服务