模式
计算机科学
背景(考古学)
话语
对抗制
人工智能
光学(聚焦)
多模式学习
利用
自然语言处理
多模态
深度学习
古生物学
社会科学
物理
计算机安全
社会学
万维网
光学
生物
作者
Minjie Ren,Xiangdong Huang,Jing Liu,Ming Liu,Xuanya Li,An-An Liu
标识
DOI:10.1109/tcsvt.2023.3273577
摘要
Multimodal emotion recognition in conversations (ERC) aims to identify the emotional state of constituent utterances expressed by multiple speakers in dialogue from multimodal data. Existing multimodal ERC approaches focus on modeling the global context of the dialogue and neglect to mine the characteristic information from the corresponding utterances expressed by the same speaker. Additionally, information from different modalities exhibits commonality and diversity for emotional expression. The commonality and diversity of multimodal information are compensated for each other but not effectively exploited in previous multimodal ERC works. To tackle these issues, we propose a novel Multimodal Adversarial Learning Network (MALN). MALN first mines the speaker’s characteristics from context sequences and then incorporate them with the unimodal features. Afterward, we design a novel adversarial module AMDM to exploit both commonality and diversity from the unimodal features. Finally, AMDM fuses different modalities to generate refined utterance representations for emotion classification. Extensive experiments are conducted on two public multimodal ERC datasets, IEMOCAP and MELD. Through the experiments, MALN shows its superiority over the state-of-the-art methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI