MuSe-Topic: Multimodal Emotion-Target Sub-challenge (MuSe2020)
收藏资源简介:
[Post-challenge] MuSe-Topic of MuSe2020: Predicting 10-class domain-specific topics as the target of 3-class (low, medium, high) emotions of valence and arousal. This package includes only MuSe-Topic features (all partitions) and annotations of the training and development set (test scoring via the MuSe website). General: The purpose of the Multimodal Sentiment Analysis in Real-life media Challenge and Workshop (MuSe) is to bring together communities from different disciplines; mainly, the audio-visual emotion recognition community (signal-based), and the sentiment analysis community (symbol-based). We introduce the novel dataset MuSe-CAR that covers the range of aforementioned desiderata. MuSe-CAR is a large (>36h), multimodal dataset that has been gathered in-the-wild with the intention of further understanding Multimodal Sentiment Analysis in-the-wild, e.g., the emotional engagement that takes place during product reviews (i.e., automobile reviews) where a sentiment is linked to a topic or entity. We have designed MuSe-CAR to be of high voice and video quality, as informative video social media content, as well as everyday recording devices have improved in recent years. This enables robust learning, even with a high degree of novel, in-the-wild characteristics, for example as related to: i) Video: Shot size (a mix of closeup, medium, and long shots), face-angle (side, eye, low, high), camera motion (free, free but stable, and free but unstable, switch, e.g., zoom, fixed), reviewer visibility (full body, half-body, face only, and hands only), highly varying backgrounds, and people interacting with objects (car parts). ii) Audio: Ambient noises (car noises, music), narrator and host diarisation, diverse microphone types, and speaker locations. iii) Text: Colloquialisms, and domain-specific terms.
【赛后阶段】MuSe2020的MuSe-Topic任务:针对价效(valence)与唤醒度(arousal)的3级(低、中、高)情感分类目标,预测10类垂直领域主题。本数据包仅包含MuSe-Topic特征(全分区)以及训练集与开发集的标注信息(测试集评分需通过MuSe官方网站完成)。 【概述】真实场景媒体多模态情感分析挑战赛与研讨会(MuSe)的宗旨是汇聚不同学科领域的社群:主要包括基于信号处理的音视频情感识别社群,以及基于符号处理的情感分析社群。本次任务推出了全新数据集MuSe-CAR,其覆盖了前述的各项需求。MuSe-CAR是一个时长超36小时的大型多模态数据集,采集自真实野外场景(in-the-wild),旨在进一步探索真实场景下的多模态情感分析——例如在产品评测(即汽车评测)中,情感与主题或实体相关联时所体现出的情感参与度。 考虑到近年来信息丰富的视频社交媒体内容以及日常录制设备的画质与音质均有提升,我们将MuSe-CAR设计为具备高音视频质量的数据集。这使得模型能够在具备大量新颖真实场景特性的情况下实现鲁棒性学习,相关特性示例包括: 一、视频维度:镜头景别(包含特写、中景、远景的混合)、面部拍摄角度(侧面、眼部特写、低角度、高角度)、摄像机运动模式(自由运动、自由且稳定、自由但不稳定、切换模式如变焦、固定)、拍摄者出镜范围(全身、半身、仅面部、仅手部)、高度多样化的背景,以及测试者与实体(汽车部件)的互动场景; 二、音频维度:环境噪音(汽车噪音、背景音乐)、旁白与主持人语音分段标注、多样化的麦克风类型,以及不同的声源位置; 三、文本维度:口语化表达与垂直领域专业术语。




