Transfer function measurements for simulating environmental noise at hearable microphones
收藏资源简介:
This dataset is supplementary material to the conference paper "Multi-Microphone Noise Data Augmentation for DNN-based Own Voice Reconstruction for Hearables in Noisy Environments" presented at ICASSP 2024 [1]. The dataset consists of impulse response measurements for 18 device users (5 female, 13 male) wearing hearable devices in both ears. The dataset was recorded in a sound-proof listening room using the Hearpiece prototype device (closed vent variant) [2] with a sampling frequency of 44.1 kHz.Impulse responses were measured with exponential sweeps from 80 Hz to 22.05 kHz with a duration of 3s played from 8 loudspeakers arranged in a circle of approximately 1.5m radius. The loudspeakers were located in the horizontal plane around the device users in 45°-steps (azimuth), starting from 22.5° to the right (where 0° is the front from the device users' perspective). The measurements are contained in the folder measurements. Each subfolder contains measurements from a different device user (e.g., VP_01). Each file contains the measurement for one direction, e.g. VP_01/data_0.npz contains the measurement of device user VP_01 for 22.5° azimuth, VP_01/data_1.npz is the measurement for the same device user for 22.5°+45° and so on.Measurements of device users where the device could not be inserted, or where the fit did not provide sufficient attenuation of external sounds to the in-ear microphone, were excluded. The impulse responses for two Hearpiece devices (closed vent), the concha and in-ear microphones were measured.A DPA 6060 lavalier clip microphone and a Tbone SC140 cardiod microphone were also included in the measurement as reference channels. The channels of the measurements (counting from 0): 0: Lavalier-microphone clipped to the shirt neck, shirt collar etc. of the device user 1: Reference microphone about 50 cm in front of the device user 2: Left in-ear microphone Hearpiece 3: Left concha microphone Hearpiece 4: Right in-ear microphone Hearpiece 5: Right concha microphone Hearpiece The measurement consists of impulse responses from the loudspeaker to the hearable device microphones and reference microphones, and corresponding transfer functions. Measurement metadata is included as well. The measurement files contain a python dictionary with the following fields: test_signal: the signal used for playback, consisting of a pause, the sweep, and another pause rec_signal: the recorded signal (sweep played from the loudspeaker, recorded at the microphones) sweep: the generated exponential sweep signal without pauses T: actual duration of the sweep (~2 Seconds) sweep_inv: inverse sweep (inverse w.r.t convolution of the sweep with the system response) sweep_inv_spectrum: spectrum of the inverse sweep f11: the frequency (in Hz) corresponding to the RampLen of the fade-in at the beginning of the sweep T_desd: desired duration of the sweep in seconds (2 Seconds) T_rec: recording duration in seconds (3 Seconds) start_frequency: Minimum frequency in the measurement / first frequency in the sweep (80 Hz) RampLen: Length of the fade-in ramp applied to the beginning of the sweep (based on a Hanning window) (2048 Samples) pre_pause_len: pause time between starting the measurement and sweep playback (88200 Samples) after_pause_len: pause time after sweep playback (44100 Samples) n_repetitions: Number of repetitions for the measurement (1) n_channels: Number of recorded channels including loopback (7 = 4 Hearpiece, 2 reference, 1 loopback) coh_mat: Mean Squared Coherence per channel (between the measured sweep and the playback sweep signal), has shape (frequencies up to samplerate/2 x channels) ir_loopback: the measured impulse response of the loopback channel, used to measure and compensate system delay from audio interface ir_mic: the measured impulse responses of the hearable and reference microphones, with shape (samples, channels) tf_mic: the measured transfer functions between the loudspeaker and the hearable and reference microphones, with shape (frequencies up to samplerate/2, channels) system_delay: the measured system delay from the audio interface (position of the peak of the correlation between playback sweep and loopback sweep signals) samplerate: The sampling rate used for the measurements (44100 Hz) This dataset is compatible with the German own voice recordings available at https://zenodo.org/records/10844599 (same participants+device insertion and measurement setup). The example script generate_indiv_noise_dataset.py can be used to augment a single-channel noise dataset to obtain simulated individual hearable noise signals,similar to [1] but using impulse responses directly as filters instead of first computing relative transfer functions and then applying them in the STFT domain. [1] M. Ohlenbusch, C. Rollwage, S. Doclo: "Multi-microphone Noise Data Augmentation for DNN-based Own Voice Reconstruction for Hearables in Noisy Environments". In: Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Seoul, South Korea, Apr. 2024, pp. 416-420.[2] F. Denk, M. Lettau, H. Schepker, S. Doclo, R. Roden, M. Blau, J.-H. Bach, J. Wellmann, and B. Kollmeier: "A One-Size-Fits-All Earpiece with Multiple Microphones and Drivers for Hearing Device Research". In: Proc. AES International Conference on Headphone Technology. San Francisco, USA, Aug. 2019.
本数据集为发表于ICASSP 2024的会议论文《面向嘈杂环境下可穿戴设备自身语音重建的基于深度神经网络(Deep Neural Network, DNN)的多麦克风噪声数据增强方法》[1]的补充材料。 该数据集包含18名设备使用者(5名女性,13名男性)双耳佩戴可穿戴设备时的脉冲响应测量数据。实验在隔音听音室内开展,采用Hearpiece原型设备(闭孔变体)[2],采样率设置为44.1 kHz。 脉冲响应通过指数扫频信号进行测量:扫频范围为80 Hz至22.05 kHz,单段扫频时长3秒,由8个扬声器播放。扬声器沿半径约1.5米的圆周等间隔排布于使用者所在的水平面内,方位角以45°为步长,从使用者正右方22.5°处起始(0°对应使用者正前方视角)。 所有测量数据存储于measurements文件夹中。每个子文件夹对应一名独立的设备使用者(例如VP_01)。每个文件包含单个方位角的测量数据:如VP_01/data_0.npz对应使用者VP_01在22.5°方位角的测量结果,VP_01/data_1.npz对应同一使用者在22.5°+45°方位角的测量结果,以此类推。对于设备无法佩戴,或佩戴后未能为入耳式麦克风提供足够外部声音衰减的使用者,其测量数据已被排除。 本次测量涵盖两台Hearpiece设备(闭孔版本)的耳甲腔与入耳麦克风。同时引入DPA 6060领夹式麦克风与Tbone SC140心形麦克风作为参考通道。 测量通道(从0开始计数): 0:夹于设备使用者衬衫领口、颈侧等位置的领夹式麦克风 1:位于设备使用者正前方约50 cm处的参考麦克风 2:Hearpiece设备左侧入耳麦克风 3:Hearpiece设备左侧耳甲腔麦克风 4:Hearpiece设备右侧入耳麦克风 5:Hearpiece设备右侧耳甲腔麦克风 6:环路通道(用于补偿音频接口引入的系统延迟) 本次测量包含扬声器至可穿戴设备麦克风与参考麦克风的脉冲响应、对应的传递函数,以及测量元数据。测量文件为Python字典格式,包含以下字段: - test_signal:用于播放的测试信号,由一段前置静音、扫频信号及一段后置静音组成 - rec_signal:录制得到的信号(扬声器播放的扫频信号经麦克风录制所得) - sweep:不含静音段的生成式指数扫频信号 - T:扫频实际时长(约2秒) - sweep_inv:逆扫频信号(与扫频信号经系统响应卷积后的逆信号) - sweep_inv_spectrum:逆扫频信号的频谱 - f11:与扫频起始段淡入斜坡长度对应的频率(单位:Hz) - T_desd:扫频的目标时长(单位:秒,2秒) - T_rec:录制总时长(单位:秒,3秒) - start_frequency:本次测量的最低频率/扫频信号的起始频率(80 Hz) - RampLen:扫频信号起始段应用的汉宁窗淡入斜坡长度(2048个采样点) - pre_pause_len:测量启动至扫频播放之间的前置静音时长(88200个采样点) - after_pause_len:扫频播放结束后的后置静音时长(44100个采样点) - n_repetitions:本次测量的重复次数(1次) - n_channels:录制的通道总数(含环路通道,共7通道:4个可穿戴设备麦克风、2个参考麦克风、1个环路通道) - coh_mat:各通道的均方相干性(实测扫频信号与播放扫频信号之间的相干性,形状为「频率维度(采样至采样率一半) × 通道维度」) - ir_loopback:环路通道的实测脉冲响应,用于测量并补偿音频接口引入的系统延迟 - ir_mic:可穿戴设备麦克风与参考麦克风的实测脉冲响应,形状为「采样点数 × 通道数」 - tf_mic:扬声器至可穿戴设备麦克风与参考麦克风的实测传递函数,形状为「频率维度(采样至采样率一半) × 通道数」 - system_delay:音频接口带来的实测系统延迟(播放扫频信号与环路录制扫频信号之间互相关峰值的位置) - samplerate:本次测量采用的采样率(44100 Hz) 本数据集与Zenodo平台上的德语自身语音录制数据集(https://zenodo.org/records/10844599)兼容,二者采用了相同的受试者群体、设备佩戴方式与测量设置。 示例脚本generate_indiv_noise_dataset.py可用于对单通道噪声数据集进行数据增强,以生成模拟的个性化可穿戴设备噪声信号。其思路与文献[1]类似,但直接将脉冲响应作为滤波器使用,而非先计算相对传递函数再在短时傅里叶变换(Short-Time Fourier Transform, STFT)域中应用。 [1] M. Ohlenbusch, C. Rollwage, S. Doclo: 《面向嘈杂环境下可穿戴设备自身语音重建的基于深度神经网络(Deep Neural Network, DNN)的多麦克风噪声数据增强方法》. 见: IEEE国际声学、语音与信号处理会议(ICASSP)论文集. 韩国首尔, 2024年4月, 第416-420页. [2] F. Denk, M. Lettau, H. Schepker, S. Doclo, R. Roden, M. Blau, J.-H. Bach, J. Wellmann, B. Kollmeier: 《一款适用于听力设备研究的多麦克风多驱动器通用尺寸耳塞》. 见: AES国际头戴式耳机技术会议论文集. 美国旧金山, 2019年8月.



