计算机科学
人工智能
语音识别
说话人识别
特征(语言学)
集合(抽象数据类型)
说话人验证
自然语言处理
钥匙(锁)
噪音(视频)
作者
Yishuang Li,Yi Yu,Yi Tu,Shuhao Deng,Weihao Gan
标识
DOI:10.1109/icassp55912.2026.11464967
摘要
In this paper, we propose a novel speaker verification (SV) method that integrates speech pre-trained models (PTMs) with a Mixture-of-Experts (MoE) mechanism. The method first extracts multilayer representations from the PTMs and inserts an MoE module after each Transformer layer. Through dynamic routing, the MoE module performs input-adaptive feature transformations, which in turn enables implicit feature selection under varying domain conditions. We then design a cascaded residual gating mechanism to propagate key shallow-layer acoustic cues and prevent their attenuation in deeper representations. Furthermore, we introduce a cross-attention fusion module that dynamically integrates multilayer features along the layer dimension, thereby jointly maintaining the robustness of shallow features and the discriminability of deep features. The fused representations are finally fed into a downstream SV network to generate highly discriminative speaker embedding. Experiments on the VoxCeleb1-O test set show that, with the Base and Large versions of PTMs, our method achieves equal error rates of 0.628% and 0.487%, respectively, which demonstrates the effectiveness and superiority of the proposed method.
科研通智能强力驱动
Strongly Powered by AbleSci AI