DFCRN: deep-learning-based audio denoising for bird monitoring
计算机科学
降噪
深度学习
语音识别
人工智能
作者
Jinlong Xu,H. Zhang,Lin Han
标识
DOI:10.1117/12.3054133
摘要
As people's awareness of ecological protection increases, bird sound monitoring has received more and more attention. Among them, using bird sound monitoring as part of audio recognition has become a hot research topic. Since bird sounds are usually collected in natural environments, they contain a lot of noise, which will affect the monitoring results. To solve this problem, this paper designs a Convolutional Recurrent Network (CRN) that enhances feature representation along the frequency axis. This method is based on the Short-time Fourier transform (STFT) features of sound signals, focuses on the complex operation features in the time-frequency domain, and designs an Decode-Encode architecture combined with a time-frequency domain enhancement network to reduce the impact of interference information, We called this network DFCRN. Experimental results on the public datasets Birdsdata and xeno-canto-ca-nv show that compared with other denoising models, the noisy signal after DFCRN enhancement achieves the best results in SegSNR and SI-SNR, and the classification accuracy on xeno-canto-ca-nv is improved by 5%, verifying the effectiveness and robustness of this method.