Dataset for: Application of Machine Learning for the Spatial Analysis of Binaural Room Impulse Responses
收藏资源简介:
This repository contains supplementary material for the paper titled `Application of Machine Learning for the Spatial Analysis of Binaural Room Impulse Responses' Available at: dx.doi.org/10.3390/app8010105 . These programs and audio files are distributed in the hopes that they will prove useful under the Creative Commons Attribution 4.0, with no warranty; or the implied warranty of merchantability or fitness for a particular problem. Please give appropriate credit for use of the material provided in this repository back to the author. In order to use the MatLab code the Auditory Toolbox by Malcolm Slaney [1] and the Cochleagram function distributed by Bin Gao [2] are required. The python scrips require the following Python libraries to be installed: Numpy[3], SciPy[4] and Tensorflow [5]. The MatLab code was tested using MatLab R2017a on a Computer running windows 7. The python code was tested using Python 3.2.5, using an anaconda Python environment - in windows command line. -- The repository contains: Folders: 1.) neg90 - This folder contains the gaussian normalisation parameters stored as text files and the weights and biases for the trained neural network - these are all for the -90° rotation neural network. 2.) pos90 - This folder contains the gaussian normalisation parameters stored as text files and the weights and biases for the trained neural network - these are all for the +90° rotation neural network. 3.) testData - this folder contains pre-generated test data for the different binaural dummy head microphones, speaker, and signal type combinations. Python Scripts: 1.) AnalyseDoA.py - A python script that can be run to test the neural network using the pre-generated test data - running the script will allow the user to input the binaural dummy head, speaker, and signal type. The important variables generated by this script are DoA - the direction of arrival for each signal in the feature vector, and yDiff - the difference between the predicted DoA and the expected direction of arrival 2.) DirectionAnalysis.py - This python file contains a set of function that are used to define the neural network, and run it. The function called DoAPrediction takes the feature vector generated by the MatLab code as its input argument, these features will then be passed to the neural network, and the output of this function is the direction of arrival predicted by the neural network for each signal. The functions: DoAAnalysis_neg90 and DoAAnalysis_pos90 are called by the DoAPrediction function, these functions create the neural network using the NN function, import the weights and biases, and passes the feature matrix (provided as input) through the neural network - the output of these functions are the predicted direction of arrival. MatLab files: 1.) runAnalysis.m - This MatLab script analyses the dataset provided as part of this repository. Users can change the variables head ('KEMAR' or 'KU100'), signalType ('directSound' or 'reflection'), and speaker ('EquatorD5' or 'Genelec8030'). This script will produce the gaussian normalised feature vector and expected direction of arrival for all signals with the defined head, signal type, and speaker combination. These variables are then saved in .mat files so they can be imported by the python scripts. 2.) BinauralModelCochlea.m - This MatLab function analyses a given binaural signal and outputs the interaural cross-correlation, interaural level difference, interaural time difference, the cochlea output for the left and right channel and the centre frequencies of the gammatone filter band. The input variables are: IR - the signal to be analysed, N - the number of gammatone filters, freqLow - the lowest centre frequency of the gammatone filter bank (centre frequency of the first gammatone filter), and freqHigh - the highest centre frequency of the gammatone filter bank (the centre frequency of the Nth gammatone filter). This function requires Malcolm Slaney's Auditory Toolbox [1] and Bin Gao's Cochleagram function [2] in order to work. 3.) generateFeatureVector.m - This MatLab function generates a feature vector from an input binaural signal x, and a version of the signal captured after the binaural dummy head has been rotated by either +90° or -90° degree (variables xPos90 and xNeg90 respectively). If the sampling frequency (Fs) isn't 44100, the signals are resampled to be at 44100. This file also contains a function 'gaussianNormalisationTestData' which gaussian normalises the data using the mean and standard deviation calculated from the data used to train the neural networks - the mean and standard deviation values are stored in the folder GMParams in the pos90 and neg90 folders. 4.) generateTestData.m - This MatLab function analyses the included binaural dataset, it takes the input variables: head - the binaural dummy head used for the measurements either 'KEMAR' or 'KU100', speaker - the speaker used for the measurements either 'EquatorD5' or 'Genelec8030', and signalType - the type of signal being analysed either 'directSound' or 'reflection'. Text files: 1.) noLayers.txt - a text file containing the number of layers used when training the neural network - with the current version of the code the neural network contains only 1 layer. 2.) README.txt - Read me file containing information about the repository. Audio files: This repository contains 1152 binaural signals half of which are direct sounds segmented from a binaural room impulse responses and the other half are reflections segmented from binaural room impulse responses (detailed in the paper this material supports) the direct sounds are recorded at angles from 0° to 357.5° in steps of 2.5° and the reflections are recorded at angles of 1° to 358.5° in steps of 2.5°. In the paper only recordings relating to signals recorded with the Equator D5 are analysed. The combination of audio files include: 1.) 144 direct sound recordings captured with the KEMAR 45BC binaural dummy head microphone and the Equator D5 speaker 2.) 144 reflection recordings captured with the KEMAR 45BC binaural dummy head microphone and the Equator D5 speaker 3.) 144 direct sound recordings captured with the KU100 binaural dummy head microphone and the Equator D5 speaker 4.) 144 reflection recordings captured with the KU100 binaural dummy head microphone and the Equator D5 speaker 5.) 144 direct sound recordings captured with the KEMAR 45BC binaural dummy head microphone and the Genelec 8030 speaker 6.) 144 reflection recordings captured with the KEMAR 45BC binaural dummy head microphone and the Genelec 8030 speaker 7.) 144 direct sound recordings captured with the KU100 binaural dummy head microphone and the Genelec 8030 speaker 8.) 144 reflection recordings captured with the KU100 binaural dummy head microphone and the Genelec 8030 speaker The files are stored using the following file naming convention: head_Test3_speaker_signalType_000_0_Degrees.wav - where _000_0 defines the azimuth direction of arrival so for example for a direct sound measured with the KEMAR unit and the Genelec8030 at 5 degrees would be 'KEMAR_Test3_Genelec8030_directSound_005_0Degrees.wav' and for a reflection measured with the KU100 and the Equator D5 at 298.5 degrees would be 'KU100_Test3_EquatorD5_reflection_298_5Degrees.wav' -- Bibliography: [1] Slaney, M. (1998). Auditory Toolbox. Palo Alto, CA. [Online]. Available: https://engineering.purdue.edu/~malcolm/interval/1998-010/ [Accessed: Oct. 27, 2017] [2] Gao, B. (2014). Cochleagram and IS-NMF2D for Blind Source Separation. [Online] Available: http://uk.mathworks.com/matlabcentral/fileexchange/48622-cochleagram-and-is-nmf2d-for-blind-source-separation?focused=3855900&tab=function [Accessed: Oct. 27, 2017] [3] NumFocus. (n.d.). NumPy. [Online]. Available: http://www.numpy.org/ [Accessed: Oct. 27, 2017] [4] SciPy. (n.d.). SciPy. [Online]. Available: https://www.scipy.org/ [Accessed: Oct. 27, 2017] [5] Google. (n.d.). TensorFlow. [Online] Available: https://www.tensorflow.org/ [Accessed: Oct. 27, 2017] -- All code and audio produced by: Michael Lovedee-Turner, PhD candidate in Music Technology at the Audio Lab, Department of Electronic Engineering, University of York Contact: mjlt500@york.ac.uk
本仓库为论文《机器学习在双耳房间脉冲响应空间分析中的应用(Application of Machine Learning for the Spatial Analysis of Binaural Room Impulse Responses)》的补充材料,论文可通过dx.doi.org/10.3390/app8010105获取。本仓库中的程序与音频文件遵循知识共享署名4.0(Creative Commons Attribution 4.0)协议发布,不提供任何明示或默示的担保,包括但不限于适销性或特定用途适用性的担保。使用本仓库提供的材料时,请务必向原作者致谢。 使用MatLab代码需依赖Malcolm Slaney开发的听觉工具箱(Auditory Toolbox)[1]以及Bin Gao发布的耳蜗图函数(Cochleagram function)[2];Python脚本则需安装以下Python库:NumPy(Numpy)[3]、SciPy(SciPy)[4]与TensorFlow(Tensorflow)[5]。 MatLab代码已在搭载Windows 7系统的计算机上通过MatLab R2017a完成测试;Python代码则在Windows命令行的Anaconda Python环境中,基于Python 3.2.5完成测试。 --- 本仓库包含以下内容: ### 文件夹 1. **neg90**:存储用于-90°旋转神经网络的高斯归一化参数(文本文件格式)以及训练好的神经网络权重与偏置参数。 2. **pos90**:存储用于+90°旋转神经网络的高斯归一化参数(文本文件格式)以及训练好的神经网络权重与偏置参数。 3. **testData**:包含针对不同双耳假人头麦克风、扬声器与信号类型组合预生成的测试数据。 ### Python脚本 1. **AnalyseDoA.py**:可运行的Python脚本,用于基于预生成的测试数据测试神经网络。运行该脚本时,用户可输入双耳假人头、扬声器与信号类型参数。本脚本生成的关键变量包括:`DoA`(特征向量中每个信号的到达方向(direction of arrival,DoA))与`yDiff`(预测到达方向与预期到达方向的差值)。 2. **DirectionAnalysis.py**:包含用于定义与运行神经网络的一系列函数。其中`DoAPrediction`函数以MatLab代码生成的特征向量作为输入参数,将特征输入神经网络后,输出每个信号由神经网络预测得到的到达方向。`DoAPrediction`函数将调用`DoAAnalysis_neg90`与`DoAAnalysis_pos90`函数:这两个函数将通过`NN`函数创建神经网络,导入权重与偏置参数,并将输入的特征矩阵送入神经网络,最终输出预测的到达方向。 ### MatLab文件 1. **runAnalysis.m**:MatLab脚本,用于分析本仓库提供的数据集。用户可修改以下变量:`head`(可选值为`'KEMAR'`或`'KU100'`)、`signalType`(可选值为`'directSound'`或`'reflection'`)与`speaker`(可选值为`'EquatorD5'`或`'Genelec8030'`)。本脚本将为指定的假人头、信号类型与扬声器组合生成高斯归一化特征向量与预期到达方向,并将这些变量保存为`.mat`文件,以供Python脚本导入使用。 2. **BinauralModelCochlea.m**:MatLab函数,用于分析指定的双耳信号,并输出双耳互相关系数、双耳声级差、双耳时间差、左右耳耳蜗输出以及伽马通滤波器(gammatone filter)组的中心频率。其输入参数包括:`IR`(待分析的信号)、`N`(伽马通滤波器的数量)、`freqLow`(伽马通滤波器组的最低中心频率,即首个伽马通滤波器的中心频率)与`freqHigh`(伽马通滤波器组的最高中心频率,即第N个伽马通滤波器的中心频率)。本函数需依赖[1]中的听觉工具箱与[2]中的耳蜗图函数方可正常运行。 3. **generateFeatureVector.m**:MatLab函数,用于从输入的双耳信号`x`,以及经过±90°旋转的双耳假人头采集的信号`xPos90`与`xNeg90`中生成特征向量。若采样频率(Fs)不为44100 Hz,则会将信号重采样至44100 Hz。本文件还包含`gaussianNormalisationTestData`函数,该函数利用训练神经网络所用数据计算得到的均值与标准差对数据进行高斯归一化,其中均值与标准差存储于`pos90`与`neg90`文件夹内的`GMParams`子文件夹中。 4. **generateTestData.m**:MatLab函数,用于分析本仓库包含的双耳数据集。其输入参数包括:`head`(测量所用的双耳假人头,可选值为`'KEMAR'`或`'KU100'`)、`speaker`(测量所用的扬声器,可选值为`'EquatorD5'`或`'Genelec8030'`)与`signalType`(待分析的信号类型,可选值为`'directSound'`或`'reflection'`)。 ### 文本文件 1. **noLayers.txt**:文本文件,存储训练神经网络时所用的层数,当前版本代码中的神经网络仅包含1层。 2. **README.txt**:说明文件,包含本仓库的相关信息。 ### 音频文件 本仓库包含1152个双耳信号,其中一半为从双耳房间脉冲响应(binaural room impulse responses)中分割得到的直达声,另一半为从双耳房间脉冲响应中分割得到的反射声(详见本仓库支撑的论文)。直达声的录制方位角(azimuth)范围为0°至357.5°,步长为2.5°;反射声的录制方位角范围为1°至358.5°,步长为2.5°。本论文仅分析了使用Equator D5扬声器录制的音频数据。 音频文件包含以下8组组合: 1. 使用KEMAR 45BC双耳假人头麦克风与Equator D5扬声器录制的144条直达声录音 2. 使用KEMAR 45BC双耳假人头麦克风与Equator D5扬声器录制的144条反射声录音 3. 使用KU100双耳假人头麦克风与Equator D5扬声器录制的144条直达声录音 4. 使用KU100双耳假人头麦克风与Equator D5扬声器录制的144条反射声录音 5. 使用KEMAR 45BC双耳假人头麦克风与Genelec 8030扬声器录制的144条直达声录音 6. 使用KEMAR 45BC双耳假人头麦克风与Genelec 8030扬声器录制的144条反射声录音 7. 使用KU100双耳假人头麦克风与Genelec 8030扬声器录制的144条直达声录音 8. 使用KU100双耳假人头麦克风与Genelec 8030扬声器录制的144条反射声录音 音频文件采用以下命名规范:`head_Test3_speaker_signalType_000_0_Degrees.wav`。例如:使用KEMAR设备与Genelec8030扬声器录制的5°直达声文件名为`KEMAR_Test3_Genelec8030_directSound_005_0Degrees.wav`;使用KU100设备与Equator D5扬声器录制的298.5°反射声文件名为`KU100_Test3_EquatorD5_reflection_298_5Degrees.wav`。 --- ### 参考文献 [1] Slaney, M. (1998). Auditory Toolbox. Palo Alto, CA. [Online]. Available: https://engineering.purdue.edu/~malcolm/interval/1998-010/ [Accessed: Oct. 27, 2017] [2] Gao, B. (2014). Cochleagram and IS-NMF2D for Blind Source Separation. [Online] Available: http://uk.mathworks.com/matlabcentral/fileexchange/48622-cochleagram-and-is-nmf2d-for-blind-source-separation?focused=3855900&tab=function [Accessed: Oct. 27, 2017] [3] NumFocus. (n.d.). NumPy. [Online]. Available: http://www.numpy.org/ [Accessed: Oct. 27, 2017] [4] SciPy. (n.d.). SciPy. [Online]. Available: https://www.scipy.org/ [Accessed: Oct. 27, 2017] [5] Google. (n.d.). TensorFlow. [Online] Available: https://www.tensorflow.org/ [Accessed: Oct. 27, 2017] --- 所有代码与音频均由Michael Lovedee-Turner制作,他是约克大学电子工程系音频实验室音乐技术方向的博士研究生。联系方式:mjlt500@york.ac.uk



