Conditional Denoising Diffusion Implicit Model for Speech Enhancement
摘要
Recently, denoising diffusion probabilistic models (DDPMs) have been effective in speech enhancement. However, existing models largely follow the original diffusion training method, ignoring the trade-off between optimization goals in different training phases. This is not only responsible for the slow training of diffusion models, but also limits their performance. And most current research based on diffusion models predicts Gaussian noise added during the training process or clean data, without considering their relationship to the model. Moreover, the reverse process still does not fully exploit the predictive power of the model. In this work, we integrate the Min-SNR weighting strategy into the training process of the existing probabilistic conditional diffusion model to efficiently exploit different time steps. Based on this strategy, we propose a joint training algorithm and objective to use intermediate data generated during the training phase to assist model learning. Furthermore, we adopt a deterministic reverse process, which is quite different from the existing one, to speed up the inference and improve the quality of the enhanced speech. Experiments show that our proposed method has faster convergence and better performance than related diffusion models and other generative models.