Deep learning for speech denoising with improved Wiener approach
摘要
Denoising or enhancing a speech signal is a separation technique considered as a supervised learning problem. Generally, supervised learning algorithms are implemented using iterative optimization algorithms like backpropagation. The separation of speech sources poses a complex challenge due to the similar spectral and amplitude characteristics of speech, noise, and music. Source separation techniques rely on signal processing algorithms such as Non-Negative Matrix Factorization (NMF) or deep neural networks. Deep neural networks are trained to learn a time–frequency representation of noisy features of the target of interest, using Ideal Binary Masks (IBM) or Ideal Ratio Masks (IRM) due to their simplicity and significant impact on speech intelligibility. This enables better understanding and high-quality extraction of speech. In this article, we compare the separation results using different training targets, including IBM, binary masks, and IRM masks, as well as the spectral magnitude of the short-term Fourier transform, combined with the Wiener method along with the Decision-Directed, HRNR, and TSNR approaches.