遇见数据集

EmoMatchSpanishDB

收藏
DataCite Commons2023-06-08 更新2024-07-28 收录
官方服务:

资源简介:

These carpete contains the datasets features used and described in the research paper entitled García-Cuesta, E., Barba, A., Gachet, D. "EmoMatchSpanishDB: Study of Speech Emotion Recognition Machine Learning Models in a New Spanish Elicited Database" , Multimedia Tools and Applications, Ed. Springer, 2023 <br> In this paper we address the task of real time emotion recognition for elicited emotions. For this purpose we have created a publicly accessible dataset composed by fifty subjects expressing the emotions of anger, disgust, fear, happiness, sadness, and surprise in Spanish language. In addition, a neutral tone of each subject has been added. This article describes how this database have been created including the recording and the performed crowdsourcing perception test in order to statistically validate the emotion of each sample and remove noisy data samples. Moreover we present a baseline comparative study between different machine learning techniques in terms of accuracy, specificity, precision, and recall. Prosodic and spectral features are extracted and used for this classification purpose. We expect that this database will be useful to get new insights within this area of study. <br> The first dataset is "EmoSpanishDB" that contains a set of 13 and 140 spectral and prosodic features for a total of 3550 audios of 50 individuals reproducing the 12 sentences for the six different emotions, ’anger, disgust, fear, happiness, sadness, surprise’ (Ekman’s basic emotions]) plus neutral. <br> The second dataset is "EmoMatchSpanishDB" and contains a set of 13 and 140 spectral and prosodic features for a total of 2050 audios of 50 individuals reproducing the 12 sentences for the six different emotions, ’anger, disgust, fear, happiness, sadness, surprise’ (Ekman’s basic emotions]) plus neutral. These 2050 audios' features are a subset of EmoSpanishDB resulting of the matched audios after application of a crowdsourcing process to validate that the elicited emotion corresponds with the expressed.<br> The third dataset is "EmoMatchSpanishDB-Compare-features.zip" that contains the COMPARE features for the experiments of dependent-speaker and LOSO. These datasets have been used in the paper "EmoMatchSpanishDB: Study of Machine Learning Models in a New Spanish Elicited Dataset" and their creation, its contents, and also a set of baseline machine learning experiments and results are fully described within it. <br> The features are available under MIT license and if you want to get access to the original raw audio files for creating your own features and research purposes you can get them under CC-BY-NC completing and signing the agreement file (EMOMATCHAgreement.docx) and sending it via email to esteban.garcia@upm.es

本数据集收录了发表于Springer旗下《Multimedia Tools and Applications》2023年卷的论文《EmoMatchSpanishDB:面向新型西班牙语诱导式语音情感识别(Speech Emotion Recognition)机器学习模型的研究》(作者:García-Cuesta, E., Barba, A., Gachet, D.)中所使用并阐述的数据集特征。 本研究针对诱导式情感的实时语音情感识别任务展开探索。为此,我们构建了一套可公开获取的数据集,涵盖50名受试对象以西班牙语表达愤怒、厌恶、恐惧、快乐、悲伤与惊讶六种基础情绪的语音数据,并额外补充了每名受试的中性语调语音样本。本文完整阐述了该数据库的构建全流程,包括语音录制环节与众包(Crowdsourcing)感知验证测试,用于通过统计学方法验证每条样本的标注情感,并剔除含噪无效样本。此外,我们针对多种机器学习技术开展了基准对比实验,评估指标涵盖准确率、特异性、精确率与召回率。本次分类任务提取并使用了韵律特征(Prosodic Features)与频谱特征(Spectral Features)。我们期望该数据库能够为该研究领域提供新的研究视角。 第一个数据集为"EmoSpanishDB",包含13维与140维的频谱及韵律特征,总计涵盖50名受试朗读12个句子所产生的3550条语音数据,对应愤怒、厌恶、恐惧、快乐、悲伤、惊讶(埃克曼基本情绪,Ekman's Basic Emotions)六种情绪及中性情绪。 第二个数据集为"EmoMatchSpanishDB",同样包含13维与140维的频谱及韵律特征,总计涵盖50名受试朗读12个句子所产生的2050条语音数据,对应六种埃克曼基本情绪与中性情绪。该数据集的2050条语音特征是EmoSpanishDB的子集,其对应的语音经过众包验证流程筛选得到,确保诱导产生的情感与受试实际表达的情感一致。 第三个数据集为"EmoMatchSpanishDB-Compare-features.zip",包含适用于依赖发话人(Dependent-speaker)实验与留一法交叉验证(LOSO, Leave-One-Subject-Out)的COMPARE特征。上述数据集均已应用于论文《EmoMatchSpanishDB:面向新型西班牙语诱导式数据集的机器学习模型研究》中,该论文完整阐述了数据集的构建流程、内容构成,以及一系列基准机器学习实验与对应结果。 本数据集的特征以MIT许可证协议发布。若您需要获取原始语音音频文件以自行提取特征并开展研究,可通过填写并签署协议文件(EMOMATCHAgreement.docx)并发送至邮箱esteban.garcia@upm.es,以CC-BY-NC许可证协议获取原始音频。

提供机构:
figshare
创建时间:
2021-03-15
搜集汇总
数据集介绍
EmoMatchSpanishDB 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务