计算机科学
人工智能
情绪分析
模式
自然语言处理
缺少数据
模态(人机交互)
机器学习
光学(聚焦)
语音识别
作者
Ziqi Shu,Rongzhou Zhou,Xiaodong Wang,Qingfeng Wu,Lu Cao
标识
DOI:10.1109/icassp55912.2026.11460508
摘要
In multimodal emotion recognition tasks, the widespread issue of missing modalities severely hinders model performance and generalization ability. To address this challenge, we propose PMoE, a Prompt-guided Mixture-of-Experts framework for robust multi-modal emotion recognition. Built upon a frozen, pretrained Transformer backbone, PMoE introduces a missing modality generation scheme that combines generative prompts and confidence-weighted fusion, effectively enhancing the quality of missing information compensation. A two-stage dynamic routing mechanism is further employed within the MoE layer to enable more flexible cross-modal feature fusion. In addition, a self-distillation strategy is adopted to stabilize training and improve generalization by leveraging historical model outputs as soft targets for progressive optimization. Experimental results on four public datasets—CMU-MOSI, CMU-MOSEI, IEMOCAP, and CH-SIMS—demonstrate that PMoE consistently outperforms existing baselines, especially under conditions of severe modality incompleteness, validating the effectiveness and robustness of the proposed framework.
科研通智能强力驱动
Strongly Powered by AbleSci AI