CAMEO
收藏资源简介:
CAMEO是一个多语言情感语音数据集的精选集合,旨在促进情感识别和其他语音相关任务的研究。该数据集包含13个经过精心挑选的数据集,涵盖8种语言,总计41,265个音频样本。数据集包含了17种独特的情感状态,其中93.45%被标注为7种主要情感:愤怒、厌恶、恐惧、幸福、中性、悲伤和惊讶。数据集的创建过程包括数据获取、数据标准化、元数据序列化和分发与文档化。CAMEO数据集旨在解决现有数据集缺乏标准化和跨语言鲁棒性的问题,为研究人员提供一个全面且可重复的基准,以评估和比较跨语言模型的性能。
CAMEO is a curated collection of multilingual emotional speech datasets intended to advance research in emotion recognition and other speech-related tasks. This collection comprises 13 carefully selected datasets covering 8 languages, with a total of 41,265 audio samples. It encompasses 17 distinct emotional states, 93.45% of which are annotated with 7 primary emotions: anger, disgust, fear, happiness, neutral, sadness, and surprise. The development workflow of the CAMEO dataset includes data acquisition, data standardization, metadata serialization, distribution and documentation. The CAMEO dataset aims to address the limitations of existing datasets in terms of standardization and cross-lingual robustness, providing researchers with a comprehensive and reproducible benchmark for evaluating and comparing the performance of cross-lingual models.




