模式
计算机科学
情绪识别
语音识别
自然语言处理
模式治疗法
人工智能
心理学
社会科学
社会学
心理治疗师
标识
DOI:10.1109/icedcs64328.2024.00019
摘要
This research aims to explore and optimize multimodal emotion recognition to enhance its performance. Multimodal emotion recognition involves analyzing information from different modalities—speech, vision, and text—to identify and classify emotional states accurately. This study investigates the roles of speech and text modalities and the auxiliary effect of visual modalities in multimodal emotion recognition. Study employs a two-stage multimodal information fusion neural network based on graph and attention mechanisms. The core idea is to achieve an effective fusion of multimodal information through graph convolutional networks and cross-modal attention mechanisms to improve emotion recognition performance. This study uses two datasets, Interactive Emotional Dyadic Motion Capture (IEMOCAP) and Multimodal Emotion Lines Dataset (MELD), to analyze the impact of speech and text modalities on visual modalities and their potential side effects. Results show that while the visual modality alone is less effective, adding speech or text significantly improves performance. However, introducing the speech modality and the text modality is less beneficial. Experimental outcomes reveal that the speech modality's effect is not as significant as the text modality, and its inclusion does not always lead to positive results, sometimes even degrading overall recognition performance.
科研通智能强力驱动
Strongly Powered by AbleSci AI