计算机科学
人工智能
状态空间
萧条(经济学)
模式识别(心理学)
语音识别
数学
统计
宏观经济学
经济
作者
Jingyi Liu,Yuanyuan Shang,Mengyuan Yang,Zhuhong Shao,Jiaxi Lu,Tie Liu
出处
期刊:
日期:2025-03-12
卷期号:: 1-5
被引量:4
标识
DOI:10.1109/icassp49660.2025.10888951
摘要
Depression is a severe mental illness, and extracting emotional information from video-audio signals for multimodal depression recognition is a challenging problem. Recent methods use the self-attention (SA) mechanism from Transformers to capture the dynamic relationships between different modalities. However, the quadratic computational complexity of SA reduces its effectiveness in modeling long sequences, making it insufficient for capturing complex intra-modal and inter-modal complementarity. To address this issue, this work proposes a Multimodal Fusion Mamba (MFMamba) framework, which is attention-free and purely focuses on using state space models (SSMs) for long-sequence modeling. Specifically, we devise Video Spatio-Temporal Mamba (VSTMamba) and Audio Temporal Mamba (ATMamba) for video-audio feature extraction. To fully capture the correlations among multimodal features and eliminate information redundancy, we introduce Fusion Mamba (FMamba) to integrate various features effectively. In experiments on AVEC 2013 and AVEC 2014 datasets, our method achieved competitive results.
科研通智能强力驱动
Strongly Powered by AbleSci AI