计算机科学
语音识别
字错误率
语音增强
初始化
背景(考古学)
噪音(视频)
词(群论)
降噪
人工智能
数学
古生物学
几何学
图像(数学)
生物
程序设计语言
作者
Heitor R. Guimarães,Hitoshi Nagano,Diego W. Silva
标识
DOI:10.1016/j.eswa.2020.113582
摘要
In this paper, we present Speech Enhancement through Wave-U-Net (SEWUNet), an end-to-end approach to reduce noise from speech signals. This background context is detrimental to several downstream systems, including automatic speech recognition (ASR) and word spotting, which in turn can negatively impact end-user applications. We show that our proposal does improve signal-to-noise ratio (SNR) and word error rate (WER) compared with existing mechanisms in the literature. In the experiments, network input is a 16 kHz sample rate audio waveform corrupted by an additive noise. Our method is based on the Wave-U-Net architecture with some adaptations to our problem. Four simple enhancements are proposed and tested with ablation studies to prove their validity. In particular, we highlight the weight initialization through an autoencoder before training for the main denoising task, which leads to a more efficient use of training time and a higher performance. Through quantitative metrics, we show that our method is prefered over the classical Wiener filtering and shows a better performance than other state-of-the-art proposals.
科研通智能强力驱动
Strongly Powered by AbleSci AI