FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition
收藏资源简介:
MDER-MA数据集是一个多模态情感识别数据集,专门针对摩洛哥阿拉伯语(Darija)设计。它包含了1216个音频与文本的配对,覆盖4种情感类别:愤怒、快乐、中性和悲伤。音频文件为48 kHz的立体声.wav格式,中位持续时间约为5秒。数据集还包含了说话者的性别、年龄、会话ID等详细信息,并且数据分割是说话者不相交的,以确保模型的泛化能力。该数据集适用于文本到语音、自动语音识别和音频分类等任务。
The MDER-MA dataset is a multimodal emotion recognition dataset specifically designed for Moroccan Arabic (Darija). It includes 1216 audio↔transcript pairs covering 4 emotion classes: Angry, Happy, Neutral, and Sad. The audio files are in 48 kHz stereo .wav format with a median duration of ≈5 seconds. The dataset also provides detailed information about the speakers gender, age, session ID, etc., and the splits are speaker-disjoint to ensure model generalization. This dataset is suitable for tasks such as text-to-speech, automatic speech recognition, and audio classification.




