语音增强
计算机科学
语音识别
残余物
人工智能
噪音(视频)
降噪
背景噪声
功率(物理)
信号处理
信噪比(成像)
声学
领域(数学)
光谱密度
光谱图
作者
Xingyu Shen,Wei‐Ping Zhu,Benoit Champagne
标识
DOI:10.1109/asru65441.2025.11434652
摘要
We propose PhysMVNet, a physics-inspired end-to-end framework for multichannel speech enhancement that integrates a learnable MVDR beamformer, a Helmholtz-inspired STFT-domain regularizer, and a residual spectral mapping module. The beamformer is trained with a reconstruction loss, while the regularizer encourages local smoothness in the STFT spectrogram to improve robustness to noise and array perturbations. To mitigate spectral distortions introduced by beamforming, we incorporate a three-band residual spectral mapping network to restore fine details. Experiments on CHiME-3/4 show that PhysMVNet achieves state-of-the-art perceptual quality and intelligibility while maintaining a lightweight design suitable for realtime application. It also remains stable under extreme low-SNR conditions and array perturbations. Ablation studies confirm the contribution of each component, highlighting the benefits of physics-inspired priors in deep beamforming networks for robust, high-fidelity speech enhancement.
科研通智能强力驱动
Strongly Powered by AbleSci AI