情态动词
多模态
模态(人机交互)
感知
代表(政治)
情绪分析
透视图(图形)
机器学习
计算机科学
边距(机器学习)
人工智能
自然语言处理
模式识别(心理学)
神经科学
化学
生物
政治
万维网
高分子化学
法学
政治学
作者
Jie Wang,Yan Yang,Keyu Liu,Zhuyang Xie,Fan Zhang,Tianrui Li
标识
DOI:10.1016/j.knosys.2024.111848
摘要
Multimodal sentiment prediction poses a formidable challenge that necessitates a profound understanding of both visual and linguistic cues, as well as the intricate interactions between them. The current achievements of modern systems in this domain can plausibly be attributed to the development of sophisticated cross-modal fusion techniques. Nevertheless, such solutions often handle each modality equally, neglecting the discordant predictions arising from sentiment incongruity in unimodal sources, which may result in performance degradation in conventional extraction-fusion scenarios. In this work, we take a different route–introducing an extraction-estimation-fusion paradigm aimed at exploring more reliable multimodal representations under the supervision of unimodal sentiment prediction. To this end, we propose a Cross-modal IncongruiTy pErception NETwork, named CiteNet, for multimodal sentiment detection. In CiteNet, we initially develop a cross-modal alignment module tailored to synchronize modality-specific representations through contrastive learning. Subsequently, with a refined cross-modal integration module, CiteNet can achieve a synergistic and comprehensive multimodal representation. In addition, we explore a cross-modal incongruity learning module from an information-theoretic perspective, capable of estimating inherent sentiment disparities by analyzing modal distributions. This incongruity score is then employed as a crucial factor in the adaptive fusion of unimodal and multimodal representations, culminating in enhanced accuracy in sentiment prediction. Experimental results on two datasets demonstrate that CiteNet outperforms prior methods by a significant margin of approximately 1%–11% in accuracy.
科研通智能强力驱动
Strongly Powered by AbleSci AI