计算机科学
接头(建筑物)
语音识别
人工智能
噪音(视频)
语音活动检测
模式识别(心理学)
信号处理
计算机视觉
欺骗攻击
语音处理
特征(语言学)
背景噪声
钥匙(锁)
作者
Minjiao Yang,Kangfeng Zheng,Jujie Wang,Xi Zhang,Yaru Zhao
标识
DOI:10.1109/icassp55912.2026.11463062
摘要
Speech spoofing detection (SSD) requires capturing complex dependencies across temporal, spectral, and both short- and long-term artifacts. While Mamba-based models have shown promise in SSD, their typically simple fusion strategies limit the joint modeling of diverse artifacts. In this work, we propose a novel Tri-Attention Fusion module that progressively integrates the bidirectional dual-branch outputs of BiMamba-ST to produce highly discriminative features. Within this module, Local-domain Attention adaptively fuses bidirectional features along the channel dimension, Cross-domain Attention enables effective spectro-temporal interactions, and Global Representation Pooling unifies the stage-wise representations. We further demonstrate the module’s versatility across different frontends, including end-to-end and pre-trained systems. Experiments on ASVspoof 2019 LA, 2021 LA, 2021 DF, and In-the-Wild datasets show that the proposed countermeasure achieves performance comparable to or exceeding current state-of-the-art methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI