计算机科学
模式
人工智能
自然语言处理
情绪分析
情态动词
变压器
社会科学
量子力学
物理
社会学
电压
化学
高分子化学
作者
Meng Xu,Feifei Liang,Xiangyi Su,Cheng Fang
出处
期刊:IEEE Access
[Institute of Electrical and Electronics Engineers]
日期:2022-01-01
卷期号:10: 131671-131679
被引量:19
标识
DOI:10.1109/access.2022.3219200
摘要
Multimodal Sentiment Analysis (MSA) is an emerging research field that aims to identify the sentiment of speakers through Audio (A), Video (V), and Text (T) modalities. The major challenge is to capture joint representations that can associate and integrate information from various modalities. Most of the existing methods are prone to acquiring joint representation through the concatenation of input features. However, these methods are short of exploiting interactions fully to ensure consistency and complementarity among modalities. To solve this problem, we design a novel multimodal sentiment analysis framework named Cross-Modal Joint Representation Transformer (CMJRT), which exploits hierarchical interactions among modalities by passing joint representations from bimodality to unimodality. Specifically, we adopt cyclic translation to obtain joint representations of bimodality, where one modality is translated to the other modality forward and backward by encoder-decoders. The translation process ensures consistency between modalities. In addition, to explore complementarity among modalities, the cross-modal transformer is used to reinforce each unimodality with common information from bimodality. Extensive experiments on CMU-MOSI and CMU-MOSEI datasets demonstrate that our proposed method outperforms existing approaches.
科研通智能强力驱动
Strongly Powered by AbleSci AI