Librispeech Slakh Unmix (LSX)
收藏资源简介:
Introduction Librispeech Slakh Unmix (LSX) is a proof of concept source separation dataset for training and testing algorithms that separate a monaural audio signal using hyperbolic embeddings for hierarchical separation. The dataset is composed of artificial mixtures using audio from the librispeech (clean subset) and Slakh2100 datasets. The dataset was introduced in our paper Hyperbolic Audio Source Separation. At a Glance The size of the unzipped dataset is ~28GB Each mixture is 60-s in length and denotes the first 60 s of the bass, drums, and guitar stems of the associated Slakh2100 track. Audio is encoded as 16 bit wav files at a sampling rate of 16 kHz The data is split into training tr (1390 mixtues), validation cv (348 mixtures) and testing tt (209 mixtures) subsets The directory for each mixture contains eight wav files: mix.wav the overall mixture from the five child sources music_mix.wav the music submix containing guitar, bass, and drums speech_mix.wav the speech submix containing both male and female speech signals bass.wav original bass submix from slakh track drums.wav original drums submix from slakh track guitar.wav original guitar submix from slakh track speech_male.wav concatenated male speech utterances filling the length of the song speech_female.wav concatenated female speech utterances filling the length of the song Other Resources Pytorch code for training models along with our hyperbolic separation interface are available here Citation If you use LSX in your research, please cite our paper: @InProceedings{Petermann2023ICASSP_hyper, author = {Petermann, Darius and Wichern, Gordon and Subramanian, Aswin and {Le Roux}, Jonathan}, title = {Hyperbolic Audio Source Separation}, booktitle = {Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)}, year = 2023, month = jun } Copyright and License The LSX dataset is released under CC-BY-4.0 license. All data: Created by Mitsubishi Electric Research Laboratories (MERL), 2022-2023 SPDX-License-Identifier: CC-BY-4.0
数据集介绍 Librispeech Slakh Unmix(LSX)是一款概念验证型声源分离数据集(source separation dataset),用于训练与测试采用双曲嵌入(hyperbolic embeddings)实现分层分离(hierarchical separation)的单声道音频信号(monaural audio signal)分离算法。该数据集通过整合Librispeech(清洁子集)与Slakh2100数据集的音频素材构建人工混合音频。本数据集首次提出于我们的论文《双曲音频声源分离(Hyperbolic Audio Source Separation)》。 概览 解压后数据集总大小约为28GB。每份混合音频时长为60秒,对应关联Slakh2100音轨中贝斯、鼓与吉他声部的前60秒片段。音频以16kHz采样率编码为16位WAV文件。数据集划分为训练集tr(1390条混合音频)、验证集cv(348条混合音频)与测试集tt(209条混合音频)子集。每份混合音频的目录包含8个WAV文件: mix.wav:来自5个子声源的总混合音频; music_mix.wav:包含吉他、贝斯与鼓的音乐子混合音频; speech_mix.wav:包含男声与女声语音信号的语音子混合音频; bass.wav:源自Slakh音轨的原始贝斯声部; drums.wav:源自Slakh音轨的原始鼓声部; guitar.wav:源自Slakh音轨的原始吉他声部; speech_male.wav:拼接至歌曲完整时长的男性语音片段; speech_female.wav:拼接至歌曲完整时长的女性语音片段。 其他资源 用于训练模型的PyTorch代码与我们的双曲分离接口可在此处获取。 引用规范 若您在研究中使用LSX数据集,请引用我们的论文: @InProceedings{Petermann2023ICASSP_hyper, author = {Petermann, Darius and Wichern, Gordon and Subramanian, Aswin and {Le Roux}, Jonathan}, title = {Hyperbolic Audio Source Separation}, booktitle = {Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)}, year = 2023, month = jun } 版权与许可 LSX数据集采用CC-BY-4.0协议发布。所有数据由三菱电机研究实验室(Mitsubishi Electric Research Laboratories, MERL)于2022-2023年创建。 SPDX-License-Identifier: CC-BY-4.0



