Speech Intelligibility Versus Congruency: User Preferences of the Acoustics of Virtual Reality Game Spaces. Supplementary Data.
收藏资源简介:
Abstract 3D audio spatializers for Virtual Reality (VR) can use the acoustic properties of the surfaces of a visualised game space to calculate a matching reverb. However, this approach could lead to reverbs that impair the tasks performed in such a space, such as listening to speech-based audio. Sound designers would then have to alter the room’s acoustic properties independently of its visualisation to improve speech intelligibility, causing audio-visual incongruency. As user expectation of simulated room acoustics regarding speech intelligibility in VR has not been studied, this study asked participants to rate the congruency of reverbs and their visualisations in 6-DoF VR while listening to speech-based audio. The participants compared unaltered, matching reverbs with sound-designed, mismatching reverbs. The latter features improved D50s and reduced RT60s at the cost of lower audio-visual congruency. Results suggest participants preferred improved reverbs only when the unaltered reverbs had comparatively low D50s or excessive ringing. Otherwise, too dry or too reverberant reverbs were disliked. The range of expected RT60s depended on the surface visualisation. Differences in timbre between the reverbs may not affect preferences as strongly as shorter RT60s. Therefore, sound designers can intervene and prioritise speech intelligibility over audio-visual congruency in acoustically challenging game spaces. The article under the same name can be found in the journal Virtual Worlds. Repository information This repository contains supplementary materials to support the article. Details of the experiment process and the results of the data analysis can be found in the article. "Aggregated Experiment Data.zip" Contains the aggregated experiment data as Excel workbooks. Each workbook contains documentation explaining the workbook's purpose and data types used. "Speech stimuli.zip" Contains audio recordings of the stimuli used in the experiment captured at the listener's default position within the modelled room. Three rooms, labelled by surface material, are used in 4 versions, resulting in a total of 12 stimuli. Please be aware that the listener could move within the room and thereby adjust the level of direct sound. The audio is captured from Unity directly via a virtual audio loopback device using an RME Babyface sound card. The files are stored in the WAV format at 48 kHz@24 bits, stereo. They are recorded in binaural audio without headphone calibration. Binauralisation is achieved via Google Resonance Audio. Please use headphones to listen to them. The raw, un-reverberated, monaural speech test signal is also included. "IRs.zip" Contains the binaural 2-channel impulse responses of the modelled rooms generated by a Dirac impulse captured at the listener's default position within the modelled rooms in binaural audio. Please be aware that the listener could move within the room. Three rooms, labelled by abbreviated surface material, are used in 4 versions with three measurements each, resulting in a total of 36 measurements. The abbreviations are Blk = Concrete Blocks, Mrb = Marble, Fab = Fabric. The audio is captured from Unity directly via a virtual audio loopback device using an RME Babyface sound card. The files are stored in the WAV format at 48 kHz@24 bits, stereo. They are recorded in binaural audio without headphone calibration. Binauralisation is achieved via Google Resonance Audio. Please use headphones to listen to them. The archive also contains a TXT file documenting the whole capturing process of the IRs and a WAV file with the impulse used to create the IRs. "Room Parameters.zip" Contains CSV files describing the room acoustical parameters of the rooms measured in IRs.zip. Each room's version has its own CSV file, resulting in a total of 12 files. The three measurements appear as separate channels. Angelo Farina's Aurora tools have been used to calculate these parameters following ISO 3382-1:2009. The archive also contains a TXT file documenting the whole capturing process of the IRs. "Screenshots.zip" Contains screenshots of each experiment's stage, including all stimuli. Only one question of the questionnaires is captured here for brevity. The screenshots are captured within Unity using a non-VR camera.
摘要 用于虚拟现实(VR)的3D音频空间化器可利用可视化游戏场景表面的声学特性,计算匹配的混响效果。然而,该方法可能会生成损害该场景内任务表现的混响,例如聆听语音类音频。此时,音频设计师需独立于场景可视化调整房间的声学特性,以提升语音清晰度,这会导致音画不一致。由于目前尚未针对VR中语音清晰度相关的模拟房间声学的用户预期开展研究,本研究邀请受试者在六自由度(6-DoF)VR环境中聆听语音类音频,并对混响效果与其可视化场景的匹配度进行评分。受试者将未经修改的匹配混响与经过音频优化的不匹配混响进行对比:后者通过降低音画匹配度为代价,优化了D50参数并缩短了RT60混响时间。研究结果表明,仅当未经修改的混响的D50相对较低或存在过度振铃效应时,受试者才会偏好优化后的混响;除此之外,过于干涩或过于混响的混响均不受受试者青睐。受试者预期的RT60范围取决于场景的表面可视化效果。混响之间的音色差异对偏好的影响可能不如更短的RT60显著。因此,在声学条件严苛的游戏场景中,音频设计师可优先提升语音清晰度,而非音画匹配度。 同名文章已发表于期刊《Virtual Worlds》。 仓库信息 本仓库包含支持该论文发表的补充材料。实验流程细节与数据分析结果可参见该论文。 Aggregated Experiment Data.zip:包含以Excel工作簿形式存储的聚合实验数据。每个工作簿均附带文档,说明其用途与所用数据类型。 Speech stimuli.zip:包含实验中使用的刺激音频录音,采集自建模房间内听众的默认位置。实验共使用3个以表面材质命名的房间,每个房间有4种版本,总计12段刺激素材。请注意,受试者可在房间内移动,从而调整直达声的音量。音频通过虚拟音频环路设备从Unity中直接导出,所用设备为RME Babyface声卡。文件以WAV格式存储,参数为48 kHz@24比特,双声道,为未经过耳机校准的双耳音频,双耳化处理通过谷歌共振音频(Google Resonance Audio)实现。请使用耳机聆听该素材。本压缩包同时包含原始的、未添加混响的单声道语音测试信号。 IRs.zip:包含建模房间的双耳双声道脉冲响应,该脉冲响应由捕捉于建模房间内听众默认位置的狄拉克脉冲生成,为双耳音频。请注意,受试者可在房间内移动。实验共使用3个以表面材质缩写命名的房间,每个房间有4种版本,每种版本进行3次测量,总计36组测量数据。缩写含义如下:Blk = 混凝土砌块(Concrete Blocks),Mrb = 大理石(Marble),Fab = 织物(Fabric)。音频通过虚拟音频环路设备从Unity中直接导出,所用设备为RME Babyface声卡。文件以WAV格式存储,参数为48 kHz@24比特,双声道,为未经过耳机校准的双耳音频,双耳化处理通过谷歌共振音频(Google Resonance Audio)实现。请使用耳机聆听该素材。本压缩包同时包含一份记录脉冲响应采集全过程的TXT文件,以及用于生成该脉冲响应的脉冲WAV文件。 Room Parameters.zip:包含用于描述IRs.zip中建模房间声学参数的CSV文件。每个房间版本对应一个CSV文件,总计12份文件,3次测量结果将作为独立声道存储。本参数通过Angelo Farina的Aurora工具,依据ISO 3382-1:2009标准计算得到。本压缩包同时包含一份记录脉冲响应采集全过程的TXT文件。 Screenshots.zip:包含各实验阶段的截图,涵盖所有刺激素材。为简洁起见,此处仅展示了问卷中的一道题目。截图通过Unity内的非VR相机采集。



