Creating speech zones with self-distributing acoustic swarms (Augmented Dataset Part 1 of 2)
收藏资源简介:
Datasets used in the paper: "Creating speech zones with self-distributing acoustic swarms" This deposit contains the <strong>first</strong> part of the augmented dataset containing simulated and real world collected data. The datasets contains 18000 training mixtures of 3-5 speakers, of which 6000 are simulated using PyRoomAcoustics, 6000 are created from synchronized real world recordings in an anechoic chamber, and 6000 are created from synchronized recordings in ordinary reverberant rooms. It also includes a validation set of 500 mixtures from reverberant rooms, and a testing set of 1000 mixtures from reverberant rooms. The source sounds are various utterances from the VCTK dataset. For real world data, the utterances are played over a Rokono Bass+ Mini Speaker. The recordings are captured from an array of 7 microphones, as they are recorded by our robotic swarm as it is distributed across the table. The recorded audio in the real world has been subjected to audio compression and decompression using the Opus Codec to enable multiple simultaneous streams. You must download <strong>both</strong> the first and the second part of this dataset in order to use it properly. To uncompress the two datasets, download both and execute: ```cat *.tar.gz.* | tar xvfz -``` Please see the Readme for more information. Please see related identifiers for other datasets.
本数据集为论文《基于自分布声学集群的语音区域创建》所用数据,本次存档包含增强数据集的**第一部分**,内含模拟采集与真实世界采集的两类语音数据。该数据集包含18000组3至5位说话人的混合语音训练样本,其中6000组通过PyRoomAcoustics模拟生成,6000组来自消声室(anechoic chamber)内的同步真实录制语音,剩余6000组来自普通混响房间内的同步录制语音。此外还包含500组混响房间场景下的混合语音验证集,以及1000组混响房间场景下的混合语音测试集。该数据集的源音频均取自VCTK数据集的各类语音片段。针对真实世界采集的语音数据,我们通过Rokono Bass+ Mini扬声器播放源语音片段,采集设备为7麦克风阵列,由分布于桌面的机器人群体完成录制过程。真实场景下录制的音频已通过Opus编解码器(Opus Codec)完成音频压缩与解压缩操作,以支持多路数据流同时传输。为正常使用该数据集,需同时下载本数据集的**两部分**。解压该双卷数据集的方法为:下载全部分卷文件后执行命令:`cat *.tar.gz.* | tar xvfz -`。更多详细信息请参阅数据集自述文件(Readme),其他相关数据集可查看相关标识符。



