稳健性(进化)
计算机科学
差异(会计)
判别式
插值(计算机图形学)
噪音(视频)
算法
编码(集合论)
频道(广播)
语音识别
缩放比例
语音增强
人工智能
模式识别(心理学)
数学
降噪
几何学
生物化学
程序设计语言
化学
集合(抽象数据类型)
业务
计算机网络
会计
图像(数学)
基因
运动(物理)
作者
Zilu Guo,Qing Wang,Jun Du,Jia Pan,Qingfeng Liu,Chin‐Hui Lee
标识
DOI:10.1109/taslp.2024.3407533
摘要
In this paper, we propose a variance-preserving interpolation framework to improve diffusion models for single-channel speech enhancement (SE) and automatic speech recognition (ASR). This new variance-preserving interpolation diffusion model (VPIDM) approach requires only 25 iterative steps and obviates the need for a corrector, an essential element in the existing variance-exploding interpolation diffusion model (VEIDM). Two notable distinctions between VPIDM and VEIDM are the scaling function of the mean of state variables and the constraint imposed on the variance relative to the mean's scale. We conduct a systematic exploration of the theoretical mechanism underlying VPIDM, and develop insights regarding VPIDM's applications in SE and ASR using VPIDM as a frontend. Our proposed approach, evaluated on two distinct data sets, demonstrates VPIDM's superior performances over conventional discriminative SE algorithms. Furthermore, we assess the performance of the proposed model under varying signal-to-noise ratio (SNR) levels. The investigation reveals VPIDM's improved robustness in target noise elimination when compared to VEIDM. Furthermore, utilizing the mid-outputs of both VPIDM and VEIDM results in enhanced ASR accuracies, thereby highlighting the practical efficacy of our proposed approach. Code and audio examples are available online https://github.com/zelokuo/VPIDM .
科研通智能强力驱动
Strongly Powered by AbleSci AI