LibriWASN
收藏资源简介:
LibriWASN is a data set whose design is based on the LibriCSS data set. The main difference is that the data was recorded by distributed devices of an acoustic sensor network, randomly positioned on a meeting table. Thus, the microphone channels between the devices show a sampling rate offset. The data set with a total length of 20 hours was recorded in two acoustically different rooms. An acoustics lab with a room reverberation time of about 200ms and a lab room with about 800ms reverberation time. Nine different devices with different numbers of channels are available: Five smartphones with a single recording channel, 2 compact microphone arrays with 6 channels, 1 compact microphone array with 4 channels, and 1 circular microphone array with 8 channels. A total of 29 channels are available in the recordings. The same LibriSpeech sentences and speakers of the LibriCSS dataset were re-recorded and the directory structures of LibriCSS were kept. The data set is organized into subsets with different percentages of speech overlap (0% - 40%). LibriWASN can be used for various research purposes, e.g., as a test set for synchronization algorithms, speech separation, diarization, and meeting transcription systems in wireless acoustic ad-hoc sensor networks. Visit https://github.com/fgnt/libriwasn for tools and scripts. To cite this dataset please refer to @InProceedings{SchTgbHaeb2023, Title = {LibriWASN: A Data Set for Meeting Separation, Diarization, and Recognition with Asynchronous Recording Devices}, Author = {Joerg Schmalenstroeer and Tobias Gburrek and Reinhold Haeb-Umbach}, Booktitle = {ITG conference on Speech Communication (ITG 2023)}, Year = {2023}, Month = {Sep}, } A preview of the paper is available from here: http://arxiv.org/abs/2308.10682
LibriWASN是一款基于LibriCSS数据集开发的数据集,其核心差异在于,该数据集的数据由随机部署于会议桌的声学传感器网络(acoustic sensor network)分布式设备采集录制,因此不同设备间的麦克风信道存在采样率偏移。 该数据集总时长为20小时,在两间声学特性迥异的房间内采集录制:分别是混响时间约200毫秒的声学实验室,以及混响时间约800毫秒的普通实验室。本次采集共使用9款不同信道数量的设备:5款单录制信道的智能手机、2款6信道紧凑型麦克风阵列、1款4信道紧凑型麦克风阵列,以及1款8信道环形麦克风阵列,本次录制共计可用信道达29个。 该数据集复用了LibriCSS数据集中的LibriSpeech语句与说话人素材并进行重录,同时保留了LibriCSS的目录组织结构。数据集按照语音重叠率(0% - 40%)划分为多个子集。LibriWASN可应用于多种研究场景,例如作为无线声学自组织传感器网络(wireless acoustic ad-hoc sensor networks)中同步算法、语音分离、diarization以及会议转录系统的测试数据集。 可访问 https://github.com/fgnt/libriwasn 获取相关工具与脚本。 引用该数据集请参考如下文献: @InProceedings{SchTgbHaeb2023, Title = {LibriWASN: A Data Set for Meeting Separation, Diarization, and Recognition with Asynchronous Recording Devices}, Author = {Joerg Schmalenstroeer and Tobias Gburrek and Reinhold Haeb-Umbach}, Booktitle = {ITG conference on Speech Communication (ITG 2023)}, Year = {2023}, Month = {Sep}, } 该论文的预印本可通过以下链接获取:http://arxiv.org/abs/2308.10682



