遇见数据集

STARSS23: Sony-TAu Realistic Spatial Soundscapes 2023

收藏
Mendeley Data2024-05-17 更新2024-06-28 收录
数据链接:
官方服务:

资源简介:

DESCRIPTION: The Sony-TAu Realistic Spatial Soundscapes 2023 (STARSS23) dataset contains multichannel recordings of sound scenes in various rooms and environments, together with temporal and spatial annotations of prominent events belonging to a set of target classes. The dataset is collected in two different countries, in Tampere, Finland by the Audio Researh Group (ARG) of Tampere University (TAU), and in Tokyo, Japan by SONY, using a similar setup and annotation procedure. The dataset is delivered in two 4-channel spatial recording formats, a microphone array one (MIC), and first-order Ambisonics one (FOA). These recordings serve as the development dataset for the DCASE 2023 Sound Event Localization and Detection Task of the DCASE 2023 Challenge. The STARSS23 dataset is a continuation of the STARSS22 dataset. It extends the previous version with the following: An additional additional 2hrs 30mins of recordings in the development set, from 5 new rooms distributed in 47 new recording clips. An additional 1hr 40mins of recordings added in the evaluation set of the dataset. 360° videos spatially and temporally aligned to the audio recordings of the dataset (apart from 12 audio-only clips). Distance labels (in cm) for the spatially annotated sound events, instead of the previous azimuth and elevation only labels. Contrary to the three previous datasets of synthetic spatial sound scenes of TAU Spatial Sound Events 2019 (development/evaluation), TAU-NIGENS Spatial Sound Events 2020, and TAU-NIGENS Spatial Sound Events 2021 associated with previous iterations of the DCASE Challenge, the STARS22-23 dataset contains recordings of real sound scenes and hence it avoids some of the pitfalls of synthetic generation of scenes. Some such key properties are: annotations are based on a combination of human annotators for sound event activity and optical tracking for spatial positions, the annotated target event classes are determined by the composition of the real scenes, the density, polyphony, occurences and co-occurences of events and sound classes is not random, and it follows actions and interactions of participants in the real scenes. The first round of recordings was collected between September 2021 and January 2022. A second round of recordings was collected between November 2022 and February 2023. Collection of data from the TAU side has received funding from Google. A demo video combining the different modalities and spatial annotations can be found here. REPORT & REFERENCE: If you use this dataset you could cite this report on its design, capturing, and annotation process: Kazuki Shimada, Archontis Politis, Parthasaarathy Sudarsanam, Daniel Krause, Kengo Uchida, Sharath Adavanne, Aapo Hakala, Yuichiro Koyama, Naoya Takahashi, Shusuke Takahashi, Tuomas Virtanen, Yuki Mitsufuji (2023). STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events, found here, and Archontis Politis, Kazuki Shimada, Parthasaarathy Sudarsanam, Sharath Adavanne, Daniel Krause, Yuichiro Koyama, Naoya Takahashi, Shusuke Takahashi, Yuki Mitsufuji, Tuomas Virtanen (2022). STARSS22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events. In Proceedings of the Detection and Classification of Acoustic Scenes and Events 2022 Workshop (DCASE2022), Nancy, France. found here. AIM: The STARSS22-23 dataset is suitable for training and evaluation of machine-listening models for sound event detection (SED), general sound source localization with diverse sounds or signal-of-interest localization, and joint sound-event-localization-and-detection (SELD). Additionally, the dataset can be used for evaluation of signal processing methods that do not necessarily rely on training, such as acoustic source localization methods and multiple-source acoustic tracking. The dataset allows evaluation of the performance and robustness of the aforementioned applications for diverse types of sounds, and under diverse acoustic conditions. Specifically the STARSS23 allows evaluation of audiovisual processing methods with a spatial dimension, such as audiovisual source localization or audiovisual object recognition. SPECIFICATIONS: General: Recordings are taken in two different sites. Each recording clip is part of a recording session happening in a unique room. Groups of participants, sound making props, and scene scenarios are unique for each session (with a few exceptions). To achieve good variability and efficiency in the data, in terms of presence, density, movement, and/or spatial distribution of the sounds events, the scenes are loosely scripted. 13 target classes are identified in the recordings and strongly annotated by humans. Spatial annotations for those active events are captured by an optical tracking system. Sound events out of the target classes are considered as interference. Occurrences of up to 3 simultaneous events are fairly common, while higher numbers of overlapping events (up to 5) can occur but are rare. Volume, duration, and data split: A total of 16 unique rooms captured in the recordings, 4 in Tokyo and 12 in Tampere (development set). 70 recording clips of 30 sec ~ 5 min durations, with a total time of ~2hrs, captured in Tokyo (development dataset). 98 recording clips of 40 sec ~ 9 min durations, with a total time of ~5.5hrs, captured in Tampere (development dataset). 79 recordings clips of 40 sec ~ 7 min durations, with a total time of ~3.5hrs, captured in both sites (evaluation dataset). A training-testing split is provided for reporting results using the development dataset. 40 recordings contributed by Sony for the training split, captured in 2 rooms (dev-train-sony). 30 recordings contributed by Sony for the testing split, captured in 2 rooms (dev-test-sony). 50 recordings contributed by TAU for the training split, captured in 7 rooms (dev-train-tau). 48 recordings contributed by TAU for the testing split, captured in 5 rooms (dev-test-tau). Audio: Sampling rate: 24kHz. Bit depth: 16 bits. Two 4-channel 3-dimensional recording formats: first-order Ambisonics (FOA) and tetrahedral microphone array (MIC). Video: Video 360° format: equirectangular Video resolution: 1920x960 Video frames per second (fps): 29.97 All audio recordings are accompanied by synchronised video recordings, apart from 12 audio recordings with missing videos (fold3_room21_mix001.wav - fold3_room21_mix012.wav) More detailed information on the dataset can be found in the included README file. SOUND CLASSES: 13 target sound event classes are annotated. The classes follow loosely the Audioset ontology. 0. Female speech, woman speaking 1. Male speech, man speaking 2. Clapping 3. Telephone 4. Laughter 5. Domestic sounds 6. Walk, footsteps 7. Door, open or close 8. Music 9. Musical instrument 10. Water tap, faucet 11. Bell 12. Knock The content of some of these classes corresponds to events of a limited range of Audioset-related subclasses. For more information see the README file. EXAMPLE APPLICATION: An implementation of a trainable model performing audio-only joint SELD, trained and evaluated with this dataset is provided here. This implementation will serve as the baseline method in the DCASE 2023 Sound Event Localization and Detection Task, under the audio-only inference track. Additionally, an implementation of a trainable model performing audiovisual SELD, trained and evaluated with this dataset is provided here. This implementation will serve as the baseline method in the DCASE 2023 Sound Event Localization and Detection Task, under the audiovisual inference track. DEVELOPMENT AND EVALUATION: The current version (Version 1.1) of the dataset includes development audio/video recordings and labels and the evaluation recordings without labels, used by the participants of Task 3 of DCASE2023 Challenge to train and validate their submitted systems (development), and produce system outputs for the challenge evaluation phase. If researchers wish to compare their system against the submissions of DCASE2023 Challenge, they will have directly comparable results if they use the evaluation data as their testing set. DOWNLOAD INSTRUCTIONS: The file foa_dev.zip, correspond to audio data of the FOA recording format. The file mic_dev.zip, correspond to audio data of the MIC recording format. The file video_dev.zip contains the common videos for both audio formats. The file metadata_dev.zip contains the common metadata for both audio formats. The file foa_eval.zip corresponds to audio data of the FOA recording format for the evaluation dataset. The file mic_eval.zip corresponds to audio data of the MIC recording format for the evaluation dataset. The file video_eval.zip contains the common videos for both audio formats of the evaluation dataset. Download the zip files corresponding to the format of interest and use your favourite compression tool to unzip these zip files.

索尼-TAU 真实空间声景2023(Sony-TAu Realistic Spatial Soundscapes 2023, STARSS23)数据集包含多种房间与环境下的多声道声景录音,以及属于若干目标类别的显著事件的时空标注。该数据集采集于两个不同国家:芬兰坦佩雷的坦佩雷大学(Tampere University, TAU)音频研究组(Audio Research Group, ARG),以及日本东京的索尼(SONY),采集采用了一致的设备配置与标注流程。 数据集以两种4声道空间录音格式交付:麦克风阵列(Microphone Array, MIC)格式与一阶 Ambisonics(First-order Ambisonics, FOA)格式。该数据集作为DCASE 2023挑战赛(Detection and Classification of Acoustic Scenes and Events 2023 Challenge)声事件定位与检测任务的开发数据集,是STARSS22数据集的续作。 STARSS23在STARSS22版本基础上新增了以下内容:开发集新增2小时30分钟录音,来自47个新录音片段,分布于5个新房间;评估集新增1小时40分钟录音;新增与数据集音频录音时空对齐的360°视频(仅12个纯音频片段除外);为空间标注的声事件提供以厘米为单位的距离标签,替代此前仅包含方位角与仰角的标注。 与此前配套于DCASE挑战赛过往迭代版本的三款合成空间声景数据集——TAU 2019空间声事件(开发/评估集)、TAU-NIGENS 2020空间声事件及TAU-NIGENS 2021空间声事件——不同,STARSS22-23数据集包含真实声景录音,因此规避了合成场景生成的部分缺陷。此类关键特性包括:标注结合了人工标注的声事件活动状态与光学追踪获取的空间位置;标注的目标事件类别由真实场景的实际构成决定;事件与声类别的密度、复音性、出现频次及共现情况并非随机分布,而是贴合真实场景中参与者的行为与互动逻辑。 首轮录音采集于2021年9月至2022年1月,第二轮采集于2022年11月至2023年2月。坦佩雷大学侧的数据采集获得了谷歌(Google)的资助。包含多模态数据与空间标注的演示视频可在此处获取。 ### 报告与引用 若使用该数据集,请引用以下关于其设计、采集与标注流程的论文:Kazuki Shimada等人(2023)的《STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events》,可在此处获取;以及Archontis Politis等人(2022)的《STARSS22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events》,收录于DCASE 2022研讨会(Detection and Classification of Acoustic Scenes and Events 2022 Workshop, DCASE2022)论文集,法国南希,可在此处获取。 ### 应用目标 STARSS22-23数据集适用于训练与评估用于声事件检测(Sound Event Detection, SED)、多类声源通用定位或目标信号定位的机器学习听觉模型,以及联合声事件定位与检测(Joint Sound-Event-Localization-and-Detection, SELD)模型。此外,该数据集可用于评估无需依赖训练的信号处理方法,例如声源定位方法与多声源声学追踪方法。该数据集可用于评估上述应用在多样化声音类型与声学条件下的性能与鲁棒性。具体而言,STARSS23可用于评估带有空间维度的视听处理方法,例如视听声源定位或视听目标识别。 ### 数据集规格 #### 通用说明 录音采集于两个不同站点,每条录音片段均来自单个独立房间的一次录音会话。每个会话的参与人员、发声道具与场景设定均为专属(仅少数例外)。为在声事件的存在性、密度、运动状态及空间分布等维度实现数据的多样性与采集效率,场景采用松散脚本化设计。数据中共标注13个目标类别,并由人工完成精细标注。上述活跃事件的空间标注通过光学追踪系统获取。非目标类别的声事件将被视为干扰信号。同时出现最多3个事件的情况较为常见,同时出现5个及以上重叠事件的情况虽存在但较为罕见。 #### 时长、体量与数据划分 数据集共采集16个独立房间的录音:东京站点4个,坦佩雷站点12个,均属于开发集。东京站点采集的开发数据集包含70条录音片段,时长范围为30秒至5分钟,总时长约2小时。坦佩雷站点采集的开发数据集包含98条录音片段,时长范围为40秒至9分钟,总时长约5.5小时。两个站点共同采集的评估数据集包含79条录音片段,时长范围为40秒至7分钟,总时长约3.5小时。 开发集提供训练-测试划分以用于结果报告:索尼贡献的训练集包含40个录音片段,采集于2个房间(dev-train-sony);索尼贡献的测试集包含30个录音片段,采集于2个房间(dev-test-sony);坦佩雷大学贡献的训练集包含50个录音片段,采集于7个房间(dev-train-tau);坦佩雷大学贡献的测试集包含48个录音片段,采集于5个房间(dev-test-tau)。 #### 音频参数 采样率为24kHz,比特深度为16比特。提供两种4声道三维录音格式:一阶 Ambisonics(FOA)与四面体麦克风阵列(MIC)。 #### 视频参数 采用360°等距柱状投影格式,分辨率为1920×960,帧率为29.97 fps。除12条纯音频录音(文件名为fold3_room21_mix001.wav至fold3_room21_mix012.wav)外,所有音频录音均配有同步视频录制。数据集的更多详细信息可参见附带的README文件。 ### 声事件类别 数据中共标注13个目标声事件类别,类别划分大致遵循Audioset本体论: 0. 女性语音(女性讲话) 1. 男性语音(男性讲话) 2. 拍手声 3. 电话声响 4. 笑声 5. 居家声响 6. 行走脚步声 7. 门开关声 8. 音乐 9. 乐器声响 10. 水龙头流水声 11. 铃声 12. 敲击声 部分类别的内容对应有限范围的Audioset相关子类,更多信息可参见README文件。 ### 示例应用 本数据集配套提供了一款仅基于音频的联合SELD可训练模型的实现,使用该数据集完成训练与评估,该实现将作为DCASE 2023挑战赛声事件定位与检测任务纯音频推理赛道的基线方法。此外,本数据集还配套提供了一款基于视听的联合SELD可训练模型的实现,使用该数据集完成训练与评估,该实现将作为DCASE 2023挑战赛声事件定位与检测任务视听推理赛道的基线方法。 ### 开发与评估 当前数据集版本为V1.1,包含开发集的音视频数据与标注,以及无标注的评估集数据,供DCASE2023挑战赛任务3的参赛选手训练与验证其提交的系统(基于开发集),并生成挑战赛评估阶段的系统输出。若研究人员希望将其系统与DCASE2023挑战赛的提交结果进行对比,使用评估集作为测试集可获得直接可比的实验结果。 ### 下载说明 foa_dev.zip对应FOA格式的音频数据;mic_dev.zip对应MIC格式的音频数据;video_dev.zip包含两种音频格式共用的视频数据;metadata_dev.zip包含两种音频格式共用的元数据;foa_eval.zip对应评估集的FOA格式音频数据;mic_eval.zip对应评估集的MIC格式音频数据;video_eval.zip包含评估集两种音频格式共用的视频数据。请下载所需格式对应的压缩包,并使用任意解压工具完成解压。

创建时间:
2023-06-28
搜集汇总
数据集介绍
STARSS23: Sony-TAu Realistic Spatial Soundscapes 2023 数据集图片
背景与挑战
背景概述
STARSS23数据集是一个多通道录音数据集,包含不同房间和环境中的声音场景,以及目标类别事件的时空标注。数据集由Tampere University和SONY合作收集,提供了两种4通道空间录音格式(FOA和MIC),并包含360°视频,适用于声音事件检测、声源定位等机器学习任务。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务