计算机科学
稳健性(进化)
保险丝(电气)
语音增强
语音识别
噪音(视频)
噪声测量
限制
人工智能
融合
背景(考古学)
融合机制
模式识别(心理学)
任务(项目管理)
降噪
工程类
化学
系统工程
语言学
哲学
电气工程
机械工程
生物
基因
图像(数学)
古生物学
脂质双层融合
生物化学
作者
Yanfeng Wu,Taihao Li,Junan Zhao,Qirui Wang,Jing Xu
标识
DOI:10.1109/lsp.2023.3290832
摘要
Robust speaker verification (RSV) under noisy con- ditions is still a challenging task. Recently, some task-specific speech enhancement (SE) approaches are proposed and achieve excellent performance on RSV. However, all these works adopt only one kind of SE network and thus can not remove noise from different aspects, limiting the performance of the RSV task. In this letter, we propose a fused SE framework (FSEF) for RSV, which integrates both T-F masking-based and feature mapping- based SE networks to collect complementary information and improve the robustness against noise. Two FESF-RSV systems are constructed based on two kinds of fusion methods: score fusion and feature fusion. In addition, we present a Multi- Scale Attentive Context Aggregation Network (MSACAN) as the backbone structure in the FSEF. The MSACAN can not only extract and fuse multi-scale features adaptively but also enhance speaker characteristics against noise and interfering speakers. Experiments conducted on the noise-simulated VoxCeleb1 dataset demonstrate both the FSEF and the MSACAN can improve the performance of RSV compared to previous approaches.
科研通智能强力驱动
Strongly Powered by AbleSci AI