计算机科学
情绪识别
人工智能
图形
语音识别
自然语言处理
融合
传感器融合
任务分析
机器学习
人工神经网络
图论
多模态
人机交互
深度学习
信息融合
模式识别(心理学)
上下文模型
隐马尔可夫模型
任务(项目管理)
作者
Jie Yang,Ruinan Shi,Xu Du,Yiqian Xie,Weiwei Wang
标识
DOI:10.1109/tlt.2026.3691830
摘要
Multimodal Emotion Recognition in Conversations (MERC) under uncertain missing modalities remains a critical challenge, especially in educational contexts. Significant research efforts have been dedicated to developing novel solutions to tackle this problem. Nevertheless, existing methods are difficult to fully utilize available information for high-quality data reconstruction, and fail to learn discriminative multimodal representations for handling uncertain missing modalities. Furthermore, the generalization of such methods in educational contexts remains largely overlooked. To address these challenges, this study proposes a novel graph fusion network approach for achieving MERC tasks in education with uncertain missing modalities. Our method first reconstructs missing modalities by using adaptive fusion weights to aggregate information from the most similar neighbors in the available modalities. The reconstructed data is processed via a mask-aware graph structure to explicitly exploit masked cues to enhance the robustness of intra-modal representations. Initial multimodal fusion representations derived from single-modal features are employed to build a dual-graph structure for further refining the multimodal representation learning of each utterance. During the training process, our method also introduces the boundary-aware contrastive learning to enable the accurate identification of samples near decision boundaries. Experiments on three multimodal datasets, including one self-collected educational dataset, CMU-MOSI, and CMU-MOSEI, demonstrates its superiority, effectiveness, and generalizability.
科研通智能强力驱动
Strongly Powered by AbleSci AI