A multi-task Generalized Speech Restoration model using Depthwise Separable Convolution Layers
摘要
Speech enhancement and restoration methods based on deep learning are crucial for improving damaged audio recordings affected by noise, reverberation, and distortion. Existing models often focus on single-task applications like denoising, super-resolution, or dereverberation, addressing only one type of speech signal degradation. However, real-world speech is often degraded by multiple distortions simultaneously, and high computational demands also pose a challenge for achieving multi-task models. To address this limitation, we propose GSR-CDNN (Generalized Speech Restoration-Depthwise Convolutional Neural Network), a novel framework that restores high-fidelity speech by handling multiple distortions simultaneously. Unlike traditional models, GSR-CDNN integrates depthwise convolution layers in both the encoder and decoder to efficiently address, reverberation, and super-resolution in a unified approach. With only 1.75 million parameters, it provides a computationally efficient solution suitable for real-world applications. Experimental results show that GSR-CDNN outperforms CMGAN in denoising (3.48 dB), dereverberation (15.10 FWSegSNR), and super-resolution (25.6 SNR on VCTK-Single), offering superior restoration across multiple distortions.