Database of Gaze, Movement and Communication Behaviour in Free Triadic Conversations
收藏资源简介:
General Information The data was collected between 16 and 28 January 2025 at the Carl von Ossietzky Universität Oldenburg, Germany. Reference: For more detailed information about the lab setup, see also: Hohmann, V., Paluch, R., Krueger, M., Meis, M. & Grimm, G. (2020). The Virtual Reality Lab: Realization and Application of Virtual Sound Environments. Ear & Hearing, 41(Supplement 1), 31S–38S. doi:10.1097/AUD.0000000000000945 Data and File Overview The data is structured in folders as follows: group/condition/variable.csv Condition Noise Type Noise Level Noise Location warmup babble noise 54 dB SPL(C) diffuse quiet no background noise 39 dB SPL(C) - SSN62 speech-shaped noise 62 dB SPL(C) diffuse SSN78 speech-shaped noise 78 dB SPL(C) diffuse babble62 babble noise 62 dB SPL(C) diffuse babble78 babble noise 77 dB SPL(C) diffuse triad babble noise, triadic conversation 68 dB SPL(C) diffuse / head level ads radio advertisements 67 dB SPL(C) top centre The noise level was measured with a measurement microphone on the table. In the 'triad' condition, the noise level was the combined level of diffuse noise (62 dB SPL) and the concurrent triadic conversation. The virtual speakers were seated in an equilateral triangle which is shifted by 60 degree relatively to the triangle formed by the participants. Overview of variables (see below for details on coordinate system and dimensions): Variable Data Size Sensor HeadPos head translation / m 30001 x 9 optical head tracking HeadRot head rotation / deg 30001 x 9 optical head tracking EOG horizontal and vertical EOG / V 30001 x 6 EOG Sensor Gaze gaze direction, global coordinate system / deg 30001 x 3 derived from HeadRot and EOG GazeLocal gaze direction, local coordinate system / deg 30001 x 3 derived from HeadRot and EOG Levels10 Leq over 10 ms / dB SPL(C) 30001 x 3 head-mounted mic. Levels50 Leq over 50 ms / dB SPL(C) 30001 x 3 head-mounted mic. LevelsTable10 Leq over 10 ms / dB SPL(C) 30001 x 1 table mic. LevelsTable50 Leq over 50 ms / dB SPL(C) 30001 x 1 table mic. VAD voice activity 30001 x 3 derived from Levels50 T sample time / s 30001 x 1 time since begin of measurement Success conversation success (*) 10 x 3 questionnaire (*) In group 4, conditions 'babble62' and 'triad', and in group 5, condition 'babble62', it is not possible to correctly assign the questionnaire data to the individuals due to missing data. Methodological Information This study presents a dataset comprising triadic face-to-face conversations conducted under various noise conditions. Head movements, electrooculography (EOG) and short-term speech levels were recorded, as was self-perceived conversation success, as measured by a questionnaire developed by Nicoras et al. (2022) [4]. Twenty-seven participants with normal hearing, aged 19–29, and fluent in German took part in this study. Participants registered in groups of three friends, with mixed genders. They were seated in an equilateral triangle with a side length of 1.5 metres around a table within a 45-channel, semi-spherical loudspeaker setup. Each participant wore an EOG sensor to measure eye blinks and movements, a head-mounted microphone to measure their speech's sound pressure level, and a head-tracking device to measure translation and rotation of the head. An additional microphone was placed in the centre of the table to measure the overall sound pressure level. Each participant was given a tablet computer to complete the questionnaires on. TASCAR was used to interface with all sensors and data streams, render background noise and log data [2]. The measurements consisted of seven conditions ('quiet', 'SSN62', 'SSN78', 'babble62', 'babble78', 'triad' and 'ads') and one training condition ('warm up'). Each condition was measured once and consisted of a five-minute-long free triadic conversation with a certain background noise. The measurement started with the training condition, after which the remaining conditions were randomised. After each condition, participants were asked to complete a questionnaire developed by Nicoras et al. (2022) [4] to assess their perceived success in the conversation. The babble noise was recorded in the canteen at the Carl von Ossietzky Universität Oldenburg, with comprehensible speech signals removed by Grimm et al. (2019) [5]. The speech-shaped noise has the frequency spectrum of a speech signal and was created from the babble noise recording. The triadic conversation, used in the condition 'triad', was scripted and recorded with background noise in order to create a Lombard effect by Gerken et al. (2020) [6]. In the 'ads' condition, radio advertisements for local businesses from the 'Mein Spot im Radio' website were used [7]. Post processing: The receiving time stamps of EOG data were de-jittered based on hardware sensor time stamps using tascar_dl_dejitter.m from the TASCAR toolbox [1, 2]. All data except the 'Success' were resampled to 100 Hz and time-aligned using tascar_dl_resample.m. Voice activity (VAD) was calculated from 'Levels50' with the function levels2vad.m from the communication behaviour toolbox [3]. For versions of these files see file gitversions. Coordinate system: The global coordinate system is a right-handed coordinate system, i.e., x is pointing to the front, y to the left, and z upwards. The origin was in the centre of the setup on floor level. The three subjects were seated at -120 degree (right), 0 degree (centre) and 120 degree (left), facing a table in the centre of the setup. Therefore their average orientation around the z-axis in global coordinates was 60 degree, 180 degree and -60 degree, respectively. Data and dimensions: HeadPos: head position in global cartesian coordinate system, for each subject x, y, z: x1, y1, z1, x2, y2, z2, x3, y3, z3 HeadRot: head rotation in euler angles in global coordinate system, for each subject Rz, Ry, Rx Rz1, Ry1, Rx1, Rz2, Ry2, Rx2, Rz3, Ry3, Rx3 EOG: for each subject EOG-horizontal Uh, EOG-vertical Uv Uh1, Uv1, Uh2, Uv2, Uh3, Uv3 All variables with 3 columns contain data for the three subjects, one column per subject. Sensors: optical head tracking: infrared based head tracking device: Qualysis Miqus M3, 6 cameras, tracking of marker crowns EOG sensor: TI ADS1115 analog-to-digital converter (res=16 Bit, fs=860 Hz), with ESP32 WiFi microcontroller board head-mounted microphone: AKG C520 table microphone: NTI Audio M2211 questionnaire: Conversation Success Questionnaire based on items by Nicoras et al. (2022) [4], except for question 5, because here triadic conversations were used. Example Scripts Gaze direction as a function of time: load('group2/quiet/T.csv'); load('group2/quiet/GazeLocal.csv'); plot( T, GazeLocal ); xlabel('experiment time / s'); ylabel('gaze direction / deg'); Gaze direction histograms: plot_gaze_histogram.m: This script plots the gaze histograms and median head positions for a given condition and group number. To generate the data plot, type plot_gaze_histogram( 2, 'quiet' ); Speech levels: plot_speech_levels.m: Create box plot of speech levels for all subjects and conditions. This function takes no options. To generate the data plot, type plot_speech_levels(); Noise levels: plot_noise_levels.m: Create a box plot of noise levels for all conditions, i.e., the median level at the table microphone position while none of the subjects was speaking. plot_noise_levels(); References [1] Grimm, G. et al., TASCAR. https://github.com/gisogrimm/tascar [2] Grimm, G., Luberadzka, J., & Hohmann V. (2019). A toolbox for rendering virtual acoustic environments in the context of audiology. Acta Acustica united with Acustica, 105(3), 566-578. https://doi.org/10.3813/AAA.919337 [3] Grimm, G., Communication Behaviour Toolbox, https://github.com/gisogrimm/communication-behaviour-toolbox [4] Nicoras, R., Buck, B., Fischer, R. L., Godfrey, M., Hadley, L. V., Smeds, K., & Naylor, G. (2025). Effective Design for Experiments on Small-Group Conversation: Insights From an Example Study. American Journal of Audiology, 34(2), 305–320. https://doi.org/10.1044/2025_AJA-24-00226 [5] Grimm, G., & Hohmann, V. (2019, December 20). First Order Ambisonics field recordings for use in virtual acoustic environments in the context of audiology. Zenodo. https://doi.org/10.5281/zenodo.3588303 [6] Gerken, M., Hendrikse, M. M. E., Hohmann, V., & Grimm, G. (2020, November 24). German Lombard conversation recordings. Zenodo. https://doi.org/10.5281/zenodo.4160499 [7] reflexmedia GmbH (2025). Mein Spot im Radio. Accessed: 24.04.2025. https://www.meinspotimradio.de/referenzen/



