Speech Dereverberation Based on Self-supervised Residual Denoising Autoencoder with Linear Decoder
摘要
A recently proposed self-supervised denoising autoencoder with linear decoder (DAELD) speech enhancement system demonstrated promising potential in the conversion of noisy speech signals to clean signals. In addition to additive noise, reverberation is another common source of distortion, caused by multi-path reflection copies of the speech signals. The characteristics of additive noise and reverberation are different, and there is an increasing demand for an effective approach that can tackle the combined effects of these two distortion sources. Based on the promising results achieved by DAELD, we propose an extension of DAELD, called the residual DAELD system (rDAELD), in this study to simultaneously perform speech dereverberation and denoising in a self-supervised learning manner. More specifically, the proposed rDAELD does not require paired training data to estimate the model parameters, thereby making this method especially suitable for real-world applications. Experimental results confirmed that rDAELD yields promising dereverberation and denoising performance under both matched and mismatched training-test conditions for simulated and measured impulse responses.