遇见数据集

MuSe-Trust: Multimodal Trustworthiness Sub-challenge (MuSe2020)

收藏
Mendeley Data2024-05-10 更新2024-06-27 收录
数据链接:
官方服务:

资源简介:

MuSe-Trust of MuSe2020: Predicting the level of trustworthiness of user-generated audio-visual content in a sequential manner utilising a diverse range of features and (optional) emotional (arousal and valence) predictions. This package includes only MuSe-Trust features (all partitions) and annotations of the training and development set (test scoring via the MuSe website). General: The purpose of the Multimodal Sentiment Analysis in Real-life media Challenge and Workshop (MuSe) is to bring together communities from different disciplines; mainly, the audio-visual emotion recognition community (signal-based), and the sentiment analysis community (symbol-based). We introduce the novel dataset MuSe-CAR that covers the range of aforementioned desiderata. MuSe-CAR is a large (>36h), multimodal dataset which has been gathered in-the-wild with the intention of further understanding Multimodal Sentiment Analysis in-the-wild, e.g., the emotional engagement that takes place during product reviews (i.e., automobile reviews) where a sentiment is linked to a topic or entity. We have designed MuSe-CAR to be of high voice and video quality, as informative video social media content, as well as everyday recording devices have improved in recent years. This enables robust learning, even with a high degree of novel, in-the-wild characteristics, for example as related to: i) Video: Shot size (a mix of closeup, medium, and long shots), face-angle (side, eye, low, high), camera motion (free, free but stable, and free but unstable, switch, e.g., zoom, fixed), reviewer visibility (full body, half-body, face only, and hands only), highly varying backgrounds, and people interacting with objects (car parts). ii) Audio: Ambient noises (car noises, music), narrator and host diarisation, diverse microphone types, and speaker locations. iii) Text: Colloquialisms, and domain-specific terms.

2020年MuSe挑战赛的MuSe-Trust任务:以序列方式预测用户生成音视频内容的可信等级,可利用多样化特征以及(可选的)情感(唤醒度(arousal)与效价(valence))预测结果。本特征包仅包含MuSe-Trust特征(全部分区)以及训练集与开发集的标注信息(测试集评分需通过MuSe官网提交)。 概述:真实场景多模态情感分析(Multimodal Sentiment Analysis)挑战赛与研讨会(Multimodal Sentiment Analysis in Real-life media Challenge and Workshop,简称MuSe)旨在汇聚不同学科的研究者群体,主要涵盖音视频情感识别领域(基于信号处理)与情感分析领域(基于符号表征)。本次任务引入了全新数据集MuSe-CAR,其覆盖了前述的各类研究需求场景。MuSe-CAR是一个时长超36小时的大型多模态数据集,采集自真实野外场景,旨在进一步推动真实场景下的多模态情感分析研究——例如汽车产品评测过程中的情感参与度,此类场景中情感与特定主题或实体直接关联。 考虑到近年来日常录制设备与信息丰富的视频社交媒体内容的音视频质量均有显著提升,我们将MuSe-CAR设计为具备高音质与高画质的数据集。这使得模型即便在面对大量真实野外场景的复杂特性时,也能实现稳健的学习,此类特性例如包括: i) 视频维度:镜头景别(特写、中景、全景混合)、人脸拍摄角度(侧面、眼部特写、低角度、高角度)、相机运动模式(自由移动、自由且稳定、自由但不稳定、变焦等切换操作、固定机位)、拍摄者出镜状态(全身、半身、仅面部、仅手部)、背景高度多样化,以及人物与物体(汽车部件)的互动行为。 ii) 音频维度:环境噪音(汽车噪音、背景音乐)、旁白与主持者语音分段(diarisation)、多样化的麦克风类型与说话者位置。 iii) 文本维度:口语化表达与领域专属术语。

创建时间:
2023-06-28
二维码
社区交流群
二维码
科研交流群
商业服务