人工智能
模式
一致性(知识库)
计算机科学
任务(项目管理)
模式识别(心理学)
构造(python库)
机器学习
模态(人机交互)
情绪识别
任务分析
多模态
噪音(视频)
模式(计算机接口)
多模式学习
接头(建筑物)
情态动词
结构化预测
模式治疗法
编码(集合论)
情感计算
多通道交互
判别式
选择(遗传算法)
特征提取
过程(计算)
作者
Qinghongya Shi,Mang Ye,Wenke Huang,Bo Du,Xiaofen Zong
标识
DOI:10.1109/tip.2025.3608664
摘要
Multimodal emotion recognition is a task that integrates textual, visual, and audio data to holistically infer an individual's emotional state. Existing research predominantly focuses on exploiting modality-specific cues for joint learning, often ignoring the differences between multiple modalities in common goal learning. Due to multimodal heterogeneity, common goal learning inadvertently introduces optimization biases and interaction noise. To address above challenges, we propose a novel approach named Gradient and Structure Consistency (GSCon). Our strategy operates at both overall and individual levels to consider balance optimization and effective interaction respectively. At the overall level, to avoid the optimization suppression of one modality on others, we construct a balanced gradient direction that aligns each modality's optimization direction, ensuring unbiased convergence. Simultaneously, at the individual level, to avoid the interaction noise caused by multimodal alignment, we align the spatial structure of samples in different modalities. The spatial structure of the samples will not differ due to modal heterogeneity, achieving effective inter-modal interaction. Extensive experiments on multimodal emotion recognition and multimodal intention understanding datasets demonstrate the effectiveness of the proposed method. Code is available at https://github.com/ShiQingHongYa/GSCon.
科研通智能强力驱动
Strongly Powered by AbleSci AI