计算机科学
情绪分析
水准点(测量)
人工智能
杠杆(统计)
自然语言处理
机器学习
监督学习
模式
模态(人机交互)
人工神经网络
社会学
社会科学
大地测量学
地理
作者
Kezhou Chen,Shuo Wang,Yanbin Hao
标识
DOI:10.1007/978-3-031-53308-2_5
摘要
Multimodal sentiment analysis (MSA) is dedicated to deciphering human emotions in videos. It is a challenging task due to the semantic disparities among various modalities (e.g., linguistic, visual, and acoustic) present in video content. To bridge these gaps, we leverage contrastive learning and introduce a novel hierarchical multimodal approach termed Hierarchical Supervised Contrastive Learning (HSCL). Initially, we utilize an unimodal fusion combined with a supervised contrastive learning strategy to distill pertinent content from each modality. Subsequently, we combine the bimodal data in pairs and further align them through supervised contrastive learning. This paired data aids in understanding the intricate nuances of multidimensional human emotions. In our method, supervised contrastive learning is tailored to accentuate the significance of label information, facilitating the extraction of sentiment cues from diverse sources. Experimental studies on two benchmark datasets demonstrate the effectiveness of our method. Specifically, compared to the baseline on the CMU-MOSI benchmark, our method achieves 2.59% accuracy improvement for emotion recognition. Codes are released at https://github.com/Turdidae810/HSCL .
科研通智能强力驱动
Strongly Powered by AbleSci AI