计算机科学
卷积(计算机科学)
语音识别
补语(音乐)
接头(建筑物)
人工智能
情绪识别
模式识别(心理学)
芯(光纤)
数据建模
核(代数)
语音处理
重叠-添加方法
钥匙(锁)
动力学(音乐)
特征提取
自然语言处理
机制(生物学)
实体造型
计算机视觉
信号处理
建筑
说话人识别
训练集
作者
G Qian,Zhenchun Lei,Sihong Liu,Changhong Liu,Aiwen Jiang
标识
DOI:10.1109/lsp.2025.3632762
摘要
Although the Conformer model excels in speech processing, its core self-attention mechanism is limited in capturing multi-scale temporal dynamics and lacks explicit modeling of frequency-domain features, both crucial for Speech Emotion Recognition (SER). To address this, we propose ConMSDMamba, a novel Conformer-based architecture for SER. Specifically, to overcome the single-scale limitation of the original self-attention, we introduce a multi-scale dilated structure with parallel dilated convolutions to capture diverse temporal contexts. We further find that combining this structure with bidirectional Mamba models long-range temporal dependencies more efficiently than multi-head self-attention. Furthermore, to complement the Conformer's time-domain focus, we design a time-frequency convolution module that incorporates a wavelet-based branch for joint time-frequency perception. Experimental results on the widely used IEMOCAP and MELD datasets demonstrate that ConMSDMamba outperforms state-of-the-art methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI