MC-EIU
收藏资源简介:
MC-EIU数据集由内蒙古大学等机构创建,是一个全面的多模态对话情感和意图联合理解数据集。该数据集包含4,970个对话视频片段,总计56,012条数据,涵盖7种情感和9种意图,支持文本、声学和视觉三种模态,以及英语和普通话两种语言。数据集的创建过程包括数据收集、预处理和多轮标注,确保数据质量和多样性。MC-EIU数据集主要应用于人机交互领域,旨在通过理解和分析对话中的情感和意图,提升机器对人类需求的理解能力和对话系统的同理心。
The MC-EIU dataset was developed by Inner Mongolia University and other institutions, and it is a comprehensive multimodal dialogue dataset for joint emotion and intention understanding. This dataset includes 4,970 dialogue video clips, totaling 56,012 data instances, covering 7 emotion categories and 9 intention categories, and supports three modalities: text, acoustic and visual, as well as two languages: English and Mandarin. The creation process of the MC-EIU dataset involves data collection, preprocessing and multi-round annotation, which ensures the data quality and diversity. The MC-EIU dataset is mainly applied in the field of human-computer interaction, aiming to improve machines' ability to understand human needs and the empathy of dialogue systems by analyzing and comprehending emotions and intentions in dialogues.
MC-EIU 数据集分析
数据集下载
- 百度云链接: 链接
- 提取码获取: 论文接受结果公布后,通过电子邮件联系作者获取。
数据集分析
数据可视化
- 图1: 情感与意图在MC-EIU数据集中的相关性可视化。每个圆圈代表特定“情感-意图”对的样本数量。较大的圆圈表示更多的样本和更高的相关性。
相关性分析
- 数据集: MC-EIU-English 和 MC-EIU-Mandarin
- 矩阵表示: 创建了两个7×9的二维矩阵,每个元素代表每个“情感-意图”对的样本数量。
- 可视化方法: 使用样本数量作为半径,在相应的矩阵位置上绘制圆圈。
观察结果
- 情感与意图的关系: 情感和意图并非严格的一对一对应关系,不同的意图对特定情感的影响不同,反之亦然。
- 例如,“Hap-Sym”与“Hap-Agr”相比,后者出现频率更高,表明“Agreeing”更可能驱动“Happy”的表达。
- 数据集差异: 英语数据集中的情感与意图的相关性比普通话数据集更为复杂。
- 例如,“Sur”情感在英语数据集中与所有意图类别相关联,而在普通话数据集中仅与6个意图类别(“Que”, “Agr”, “Con”, “Sug”, “Wis”, 和 “Neu”)相关联。
- 模型性能: 由于这种复杂关系,模型在英语数据集上的表现相对低于普通话数据集。

- 1Emotion and Intent Joint Understanding in Multimodal Conversation: A Benchmarking Dataset内蒙古大学 · 2024年



