Mexican Emotional Speech Database (MESD)
收藏资源简介:
The Mexican Emotional Speech Database (MESD) provides single-word utterances for anger, disgust, fear, happiness, neutral, and sadness affective prosodies with Mexican cultural shaping. The MESD has been uttered by both adult and child non-professional actors: 3 female, 2 male, and 6 child voices are available (female mean age ± SD = 23.33 ± 1.53, male mean age ± SD = 24 ± 1.41, and children mean age ± SD = 9.83 ± 1.17). Words for emotional and neutral utterances come from two corpora: (corpus A) composed of nouns and adjectives that are repeated across emotional prosodies and types of voice (female, male, child), and (corpus B) which consists of words controlled for age-of-acquisition, frequency of use, familiarity, concreteness, valence, arousal, and discrete emotion dimensionality ratings. The audio recordings took place in a professional studio with the following materials: (1) a Sennheiser e835 microphone with a flat frequency response (100 Hz to 10 kHz), (2) a Focusrite Scarlett 2i4 audio interface connected to the microphone with an XLR cable and to the computer, and (3) the digital audio workstation REAPER (Rapid Environment for Audio Production, Engineering, and Recording). Audio files were stored as a sequence of 24-bit with a sample rate of 48000Hz. The amplitude of acoustic waveforms was rescaled between -1 and 1. Two speaker-embedded naturalness-reduced versions were created out of human emotional utterances for female voices from corpus B. Specifically, naturalness was progressively reduced from human voices to level 1 to level 2. In particular, duration and median pitch were edited on stressed syllables to reduce the difference between stressed and unstressed syllables. On whole utterances, F2/F1 and F3/F1 ratios were lowered by editing F2 and F3 frequencies. Intensity of harmonics 1 and 4 were also reduced. 24 utterances per emotion are available for each type of voice, corpus, and level of naturalness. They are shared as audio files in WAV format. Please see README for audio files nomenclature explanation. The MESD seems to be the first set of single-word emotional utterances that includes both adult and child voices for the Mexican population. Additionally, the MESD provides naturalness-reduced versions of emotional utterances. Citation M. M. Duville, L. M. Alonso-Valerdi, and D. Ibarra-Zarate, “The Mexican Emotional Speech Database (MESD): elaboration and assessment based on machine learning,” 43rd Annual International Conference of the IEEE Engineering in Medicine and Biology Society, p. 4, 2021. Duville, M.M.; Alonso-Valerdi, L.M.; Ibarra-Zarate, D.I. Mexican Emotional Speech Database Based on Semantic, Frequency, Familiarity, Concreteness, and Cultural Shaping of Affective Prosody. Data 2021, 6, 130. https://doi.org/10.3390/data6120130
墨西哥情感语音数据库(Mexican Emotional Speech Database, MESD)收录了带有墨西哥文化特征的愤怒、厌恶、恐惧、快乐、中性及悲伤六种情感韵律的单词语音话语。 该数据库的语音话语均由成年及儿童非专业配音员录制:共包含3名女性、2名男性及6名儿童的语音样本,其中女性平均年龄±标准差(Standard Deviation, SD)=23.33±1.53岁,男性平均年龄±标准差(Standard Deviation, SD)=24±1.41岁,儿童平均年龄±标准差(Standard Deviation, SD)=9.83±1.17岁。 情感及中性语音话语所用词汇来自两个语料库:(语料库A)由在不同情感韵律及语音类型(女性、男性、儿童)中重复出现的名词与形容词构成;(语料库B)则包含经过习得年龄、使用频率、熟悉度、具体性、效价、唤醒度及离散情感维度评分控制的词汇。 音频录制在专业录音棚中完成,所用设备如下:(1)频率响应平坦(100Hz至10kHz)的森海塞尔Sennheiser e835麦克风;(2)通过XLR线缆连接麦克风并接入电脑的福克斯特Focusrite Scarlett 2i4音频接口;(3)数字音频工作站REAPER(全称Rapid Environment for Audio Production, Engineering, and Recording,即音频制作、工程与录制快速环境)。 音频文件以24位深度、48000Hz采样率的格式存储,声学波形的振幅被重新缩放至-1至1区间内。 针对语料库B中的女性语音情感话语,我们构建了两种嵌入发声者信息的自然度降低版本。具体而言,自然度从原始人类语音开始逐步降低至等级1,再降至等级2。 具体操作上,我们对重读音节的时长及基频中值进行编辑,以缩小重读音节与非重读音节之间的差异;针对整段语音话语,通过调整第二共振峰(Formant 2, F2)与第一共振峰(Formant 1, F1)、第三共振峰(Formant 3, F3)与第一共振峰(Formant 1, F1)的比值降低其数值,同时还降低了第1和第4次谐波的强度。 针对每种语音类型、语料库及自然度等级,每种情感均包含24条语音话语。所有音频文件均以WAV格式提供。有关音频文件的命名规则说明,请参阅README文档。 墨西哥情感语音数据库(MESD)或许是首个面向墨西哥人群、同时涵盖成人与儿童发声者的单词语音情感话语数据集。此外,该数据库还提供了情感语音话语的自然度降低版本。 引用信息如下:M. M. Duville、L. M. Alonso-Valerdi及D. Ibarra-Zarate,《墨西哥情感语音数据库(MESD):基于机器学习的构建与评估》,第43届IEEE工程医学与生物医学学会年度国际会议,第4页,2021年。Duville, M.M.; Alonso-Valerdi, L.M.; Ibarra-Zarate, D.I. 基于情感韵律的语义、频率、熟悉度、具体性与文化特征构建的墨西哥情感语音数据库. Data 2021, 6, 130. https://doi.org/10.3390/data6120130



