<p>Speech enhancement and restoration methods based on deep learning are crucial for improving damaged audio recordings affected by noise, reverberation, and distortion. Existing models often focus on single-task applications like denoising, super-resolution, or dereverberation, addressing only one type of speech signal degradation. However, real-world speech is often degraded by multiple distortions simultaneously, and high computational demands also pose a challenge for achieving multi-task models. To address this limitation, we propose GSR-CDNN (Generalized Speech Restoration-Depthwise Convolutional Neural Network), a novel framework that restores high-fidelity speech by handling multiple distortions simultaneously. Unlike traditional models, GSR-CDNN integrates depthwise convolution layers in both the encoder and decoder to efficiently address, reverberation, and super-resolution in a unified approach. With only 1.75&#xa0;million parameters, it provides a computationally efficient solution suitable for real-world applications. Experimental results show that GSR-CDNN outperforms CMGAN in denoising (3.48 dB), dereverberation (15.10 FWSegSNR), and super-resolution (25.6 SNR on VCTK-Single), offering superior restoration across multiple distortions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A multi-task Generalized Speech Restoration model using Depthwise Separable Convolution Layers

  • Ikram Azaz,
  • Tao Zhang,
  • Yasir Iqbal,
  • Xin Zhao,
  • Yanzhang Geng

摘要

Speech enhancement and restoration methods based on deep learning are crucial for improving damaged audio recordings affected by noise, reverberation, and distortion. Existing models often focus on single-task applications like denoising, super-resolution, or dereverberation, addressing only one type of speech signal degradation. However, real-world speech is often degraded by multiple distortions simultaneously, and high computational demands also pose a challenge for achieving multi-task models. To address this limitation, we propose GSR-CDNN (Generalized Speech Restoration-Depthwise Convolutional Neural Network), a novel framework that restores high-fidelity speech by handling multiple distortions simultaneously. Unlike traditional models, GSR-CDNN integrates depthwise convolution layers in both the encoder and decoder to efficiently address, reverberation, and super-resolution in a unified approach. With only 1.75 million parameters, it provides a computationally efficient solution suitable for real-world applications. Experimental results show that GSR-CDNN outperforms CMGAN in denoising (3.48 dB), dereverberation (15.10 FWSegSNR), and super-resolution (25.6 SNR on VCTK-Single), offering superior restoration across multiple distortions.