遇见数据集

VocalScape: A synthetic dataset of real-world audio scenarios

收藏
Zenodo2026-02-27 更新2026-05-26 收录
官方服务:

资源简介:

VocalScape: A synthetic dataset of real‑world audio scenarios VocalScape is a synthetic soundscape dataset designed for separating human speech (foreground) from environmental sounds (background) in realistic indoor and outdoor scenarios. The dataset is generated with Soundloom, a rule‑based framework that composes mixtures from heterogeneous audio sources under explicit temporal overlap constraints, loudness normalization and consistent audio parameters. Foreground signals are drawn from Mozilla Common Voice, while background signals are constructed from environmental events taken from ESC‑50 and UrbanSound8K, providing a diverse range of non‑speech soundscapes. Data contents and structure Each VocalScape instance consists of a pair of mono flaceforms at 10 kHz sample rate: one background.flac file containing the isolated environmental component and one foreground.flac file containing the isolated speech component. The mixture used in our experiments can be reconstructed as the linear sum of these two signals. The dataset is organized into numerically indexed folders (0_0, 0_1, …), where each folder corresponds to a single mixture and contains both background.flac and foreground.flac. The archive also includes three CSV files defining the fixed dataset split: train.csv, validation.csv and test.csv. Each row lists the relative paths of the background and foreground files for a given mixture, in this order: background_path,foreground_path 0_0/background.flac,0_0/foreground.flac 0_1/background.flac,0_1/foreground.flac Foreground and background source segments are assigned to train, validation or test at the corpus level before mixture generation, so that no underlying source clip appears across different splits. This setup supports rigorous training and evaluation without leakage between train and test data. Intended use The primary use case of VocalScape is supervised separation of human speech from environmental background noise in soundscape‑like conditions. Models can be trained by reconstructing the mixture from background.flac and foreground.flac and using the individual components as clean targets. Beyond source separation, the dataset is suitable for speech enhancement, robustness studies for automatic speech recognition, and experiments on sound event detection and acoustic scene analysis that exploit the structured background component. Tooling and code To cover dataset generation, interaction, and modeling on VocalScape, we provide three complementary open‑source components: Soundloom ruleset: https://github.com/matcarollo/soundloom VocalScape dataset tooling and loaders: https://github.com/matcarollo/vocalscape MDemucs model and training code: https://github.com/matcarollo/mdemucs These repositories provide end‑to‑end support: generating VocalScape from scratch, interacting with the dataset (including Zenodo integration), and training/evaluating the separation model. License and data sources The VocalScape dataset is released under the MIT License. The underlying source corpora (Mozilla Common Voice, ESC‑50, UrbanSound8K) remain governed by their original licenses and terms of use; users of VocalScape are responsible for complying with both the VocalScape license and the licenses of these underlying datasets.

提供机构:
Zenodo
创建时间:
2026-01-25
二维码
社区交流群
二维码
科研交流群
商业服务