遇见数据集

Ablation experiment results comparison.

收藏
Figshare2025-11-21 更新2026-04-28 收录
官方服务:

资源简介:

In modern multimodal interaction design, integrating information from diverse modalities—such as speech, vision, and text—presents a significant challenge. These modalities differ in structure, timing, and data volume, often leading to mismatches, low computational efficiency, and suboptimal user experiences during the integration process. This study aims to enhance both the efficiency and accuracy of multimodal information fusion. To achieve this, publicly available datasets—Carnegie Mellon University Multimodal Opinion Sentiment Intensity (CMU-MOSI) and Interactive Emotional Dyadic Motion Capture (IEMOCAP)—are employed to collect speech, visual, and textual data relevant to multimodal interaction scenarios. The data undergo preprocessing steps including noise reduction, feature extraction (e.g., Mel Frequency Cepstral Coefficients and keypoint detection), and temporal alignment. An improved Kuhn-Munkres algorithm is then proposed, extending the traditional bipartite graph matching model to support weighted multimodal matching. The algorithm dynamically adjusts weight coefficients based on the importance scores of each modality, while also incorporating a cross-modal correlation matrix as a constraint to improve the robustness of the matching process. The enhanced algorithm’s performance is validated through information matching efficiency tests and user interaction satisfaction surveys. Experimental results show that it improves multimodal information matching accuracy by 28.2% over the baseline method. Integration efficiency increases by 18.7%, and computational complexity is significantly reduced, with average computation time decreased by 15.4%. User satisfaction also improves, with a 19.5% increase in experience ratings. Ablation studies further confirm the critical contribution of both the dynamic weighting mechanism and the correlation matrix constraint to the overall performance. This study introduces a novel optimization strategy for multimodal information integration, offering substantial theoretical value and broad applicability in intelligent interaction design and human-computer collaboration. These advancements contribute meaningfully to the development of next-generation multimodal interaction systems.

在现代多模态(multimodal)交互设计领域,融合语音、视觉、文本等多模态信息是一项极具挑战性的任务。各类模态在结构、时序与数据量上存在显著差异,常导致融合过程中出现匹配失准、计算效率低下以及用户体验不佳等问题。本研究旨在提升多模态信息融合的效率与准确性。为此,本研究采用公开数据集:卡内基梅隆大学多模态观点情感强度数据集(Carnegie Mellon University Multimodal Opinion Sentiment Intensity,简称CMU-MOSI)与交互式情感双人动作捕捉数据集(Interactive Emotional Dyadic Motion Capture,简称IEMOCAP),收集多模态交互场景下的语音、视觉与文本数据。数据预处理环节包括降噪、特征提取(如梅尔频率倒谱系数(Mel Frequency Cepstral Coefficients)与关键点检测)及时序对齐。本研究提出一种改进的库恩-曼克莱斯算法(Kuhn-Munkres algorithm),将传统二分图匹配模型拓展为支持加权多模态匹配的框架。该算法可依据各模态的重要性得分动态调整权重系数,并引入跨模态相关矩阵作为约束条件,以提升匹配过程的鲁棒性。通过信息匹配效率测试与用户交互满意度调研,对改进后算法的性能进行验证。实验结果表明,相较于基准方法,该算法可将多模态信息匹配准确率提升28.2%,融合效率提升18.7%;同时计算复杂度显著降低,平均计算耗时缩短15.4%。用户满意度亦有所提升,体验评分提高19.5%。消融实验进一步证实,动态权重机制与相关矩阵约束对模型整体性能均起到关键作用。本研究为多模态信息融合提出了一种全新的优化策略,在智能交互设计与人机协作领域具备较高的理论价值与广泛的应用前景,将有力推动下一代多模态交互系统的发展。

创建时间:
2025-11-21
二维码
社区交流群
二维码
科研交流群
商业服务