EMOSET
收藏资源简介:
EMOSET是由University of Augsburg的研究人员Maurice Gerczuk、Shahin Amiriparian、Sandra Ottl和Björn W. Schuller共同创建的一个大规模情感语音数据集。该数据集整合了来自26个现有语音情感识别(SER)语料库的84,181个音频记录,总时长超过65小时。EMOSET不仅包含已发布的SER数据库,还包括一些未发布的语音数据库,这些数据来源于健康与福祉嵌入式智能主席(University of Augsburg),用于进一步增强训练数据,以提高深度学习模型的泛化能力和减少过拟合问题。数据集中的每个数据集都包括分类情感标签,总体上共有84,161个样本,总时长为65.6小时,所有音频记录的平均时长为2.81秒。
EMOSET is a large-scale emotional speech dataset created by researchers Maurice Gerczuk, Shahin Amiriparian, Sandra Ottl and Björn W. Schuller from the University of Augsburg. It compiles 84,181 audio recordings from 26 existing speech emotion recognition (SER) corpora, with a total duration exceeding 65 hours. EMOSET includes not only published SER databases, but also some unpublished speech datasets sourced from the Chair of Embedded Intelligence for Health and Wellbeing at the University of Augsburg. These data are used to further enhance the training corpus, so as to improve the generalization ability of deep learning models and mitigate overfitting issues. Each constituent dataset within the corpus is paired with categorical emotion labels. In total, the full dataset contains 84,161 samples, with a total duration of 65.6 hours, and the average duration of all audio recordings is 2.81 seconds.




