遇见数据集

The French Lombard Dataset

收藏
Zenodo2025-10-13 更新2026-05-26 收录
官方服务:

资源简介:

The French Lombard Dataset OVERVIEW Speech and Electroglottographic (EGG) signals were recorded from 38 speakers, comprising 19 male and 19 female speakers, producing Lombard speech under four different levels of white noise played through headphones (0, 65, 75, 85 dB SPL). The dataset includes approximately 8 hours of speech in total, with about 2 hours per noise condition and an average of 12 minutes per speaker. TASK AND EXPERIMENTAL CONDITIONS The experiment involved reading 60 sentences selected from the Fharvard database at four increasing levels of vocal intensity. Variations in vocal intensity were elicited by playing different levels of broadband noise through the participants' headphones : 1st level : No noise played through the headphones (silence); 2nd level : 65 dB SPL (noise1); 3rd level : 75 dB SPL (noise2); 4th level : 85 dB SPL (noise3); The experiment was divided into three sessions, each involving the recording of 20 sentences at the four intensity levels. Participants were given a 5-minute break between sessions to rest. To avoid adaptation effects, both the sentence order and noise levels were randomized, while ensuring that 20 sentences were recorded at each intensity level across each session. MATERIALS Headphones (audio-technica BPHS1, closed-back) were used to play ambient noise to participants.Noise level was calibrated using an artificial head with a conditioning amplifier (B&K Nexus 2690)Sentences to read were displayed on a screen placed 60cm in front of the subject. The audio signal was recorded with a 1/4-in pressure microphone (Bruël and Kjær 4189), placed 30 cm away from the participants" lips, then amplified (conditioning amplifier Bruël and Kjær 2690) and digitized with 16-bit resolution at a rate of 44.1 kHz using a FireWire audio interface (RME Fireface 800). Participants remained seated in the same position throughout the experiment, with their heads resting against the seat back to maintain a constant mouth-to-microphone distance. The sound pressure level can be calibrated using the internal reference signal of the conditioning amplifier (1 kHz at 1 V RMS, stored in the 'calibration' folder of this database), with a transduction ratio of 1 V/Pa. The electroglottographic signal (EGG) was simultaneously recorded with a two-channel electroglottograph (Glottal Enterprises EG2), using medical gel to improve electric contact between the skin and the electrodes. Electrodes were placed on both sides of the thyroid cartilage while the participants were producing a sustained vowel in their comfortable voice. No automatic gain control was used. The EGG signal was also digitized with 16-bit resolution at a rate of 44.1 kHz using the same FireWire audio interface (RME Fireface 800). FILE FORMAT There are four subfolders in this repository: "calibration", "raw", "process", and "txt". The "calibration" folder contains the internal reference signal of the conditioning amplifier (1 kHz at 1 V RMS), which is required to compute the sound pressure level of the recordings, given that the amplifier’s transduction ratio was set to 1 V/Pa. The "txt" folder contains the transcriptions of each participant's recordings. Each file corresponds to one participant, with each row representing a recorded sentence. The columns provide the following information: the recording order of the stimuli ; the list from the Fharvard corpus from which the sentence was taken (List) ; the sentence number within that list (Sentence) ; the corresponding text that was read (Text) ; the noise intensity level (Condition) ; and the session number (Session). The "raw" folder contains the raw data recordings, sampled at 44.1 kHz. Each subfolder corresponds to a participant, and .wav files within are numbered from 0 to 239, matching the rows in the corresponding .txt files. The "process" folder includes separate subfolders for each signal type (egg and wav), with subfolders for each participant. All recordings stored in the "process" folder have been resampled to a frequency of 16 kHz and the filenames have been updated to include metadata about the recording conditions: “NN_SN_LN_XXXX_XXXXXX”, where Characters 1-2 indicate the speaker number. For instance, "10" refers to speaker number 10. Characters 3-4 indicate the sessions ("s1" refers to Session 1, "s2" refers to Session 2 and "s3" refers to Session 3). Characters 5-6 indicate the list from the Fharvard corpus. For instance, "l6" refers to list number 6. Characters 7-10 indicate the sentence from the list of the Fharvard corpus. For instance, "sen5" refers to the sentence 5. Characters 11-16 indicate the noise condition, i.e., noise0, noise1, noise2, and noise3 for 0, 65, 75 and 85 dB SPL, respectively. The file "speakers.txt" provides socio-demographic information about the participants. STATISTICS Total Clips 9,120Total Duration 7:36:00Duration per Noise 01:54:00Duration per Speaker 00:12:00 LICENCE The FLombard dataset is released under the CC BY-NC-SA licence TERMS OF USE The data provided must not be used to generate any audio track imitating or reproducing the voice of the participants in the production of utterances that are not included in the corpus. This includes, but is not limited to, the prohibition of training voice cloning, text-to-speech, or voice conversion models that would target the identity of the participants of this corpus. CHANGELOG * 1.0 (May 28, 2025): Initial release. Due to a data embargo agreement signed by the participants, this repository contains only a subset of the full dataset. The complete dataset, comprising 38 participants, will be made publicly available in October 2025. * 1.1 (October 13, 2025): Updated version including all participants. CREDITS This dataset consists of excerpts from the following works: * Vincent Aubanel, C. Bayard, A. Strauß, J.-L. Schwartz, The Fharvard corpus: A phonemically-balanced French sentence resource for audiology and intelligibility research, Speech Communication, Volume 124, 2020, Pages 68-74 Recordings and annotation by Maxime Jacquelin. To cite this work, you can use the following:@dataset{jacquelin_2025_15533059, author = {Jacquelin, Maxime and Garnier, Maëva and Girin, Laurent and Vincent, Rémy and Perrotin, Olivier}, title = {The French Lombard Dataset}, month = may, year = 2025, publisher = {Zenodo}, doi = {10.5281/zenodo.15533059}, url = {https://doi.org/10.5281/zenodo.15533059},}

提供机构:
Zenodo
创建时间:
2025-10-13
二维码
社区交流群
二维码
科研交流群
商业服务