<p>Real-time monaural speech enhancement (SE) demands compact, efficient models that maintain speech clarity and intelligibility on resource-limited devices. Many deep models struggle with computational efficiency, latency, and preserving speech quality, especially using recurrent layers for temporal dependencies. This paper proposes the Complex Depthwise Convolutional Recurrent Network (CDCRN), optimized for such environments. CDCRN integrates depthwise separable convolution layers in the encoder-decoder modules and Long Short-Term Memory (LSTM) layers, reducing complexity while preserving speech quality. To further boost SE performance, self-attention layers are applied after each LSTM layers, enhancing temporal relationships by focusing on complex interactions between the real and imaginary components of spectral features, improving speech-noise separation. The model enhances speech signals directly in the time domain by minimizing the Signal to Distortion Ratio (SDR) loss, comparing predicted outputs with target clean speech. Internal layers capture complex spectral features, enabling effective noise separation and improved speech quality. Experimental results show that CDCRN, with only 1.409&#xa0;M parameters, outperforms recent LSTM-based models and baselines, achieving superior speech quality and intelligibility. Specifically on the WSJ0 dataset, CDCRN delivers a 25.08% improvement in STOI (93.33%) and a 1.52-point increase in PESQ (3.25), with strong results under seen and unseen noise conditions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speech enhancement using complex depthwise convolutional recurrent network with self-attention mechanism

  • Yasir Iqbal,
  • Tao Zhang,
  • Anjum Iqbal,
  • Jiajia Yan,
  • Ikram Azaz,
  • Qiyun Zhang,
  • Ijaz Hussain,
  • Xin Zhao,
  • Yanzhang Geng

摘要

Real-time monaural speech enhancement (SE) demands compact, efficient models that maintain speech clarity and intelligibility on resource-limited devices. Many deep models struggle with computational efficiency, latency, and preserving speech quality, especially using recurrent layers for temporal dependencies. This paper proposes the Complex Depthwise Convolutional Recurrent Network (CDCRN), optimized for such environments. CDCRN integrates depthwise separable convolution layers in the encoder-decoder modules and Long Short-Term Memory (LSTM) layers, reducing complexity while preserving speech quality. To further boost SE performance, self-attention layers are applied after each LSTM layers, enhancing temporal relationships by focusing on complex interactions between the real and imaginary components of spectral features, improving speech-noise separation. The model enhances speech signals directly in the time domain by minimizing the Signal to Distortion Ratio (SDR) loss, comparing predicted outputs with target clean speech. Internal layers capture complex spectral features, enabling effective noise separation and improved speech quality. Experimental results show that CDCRN, with only 1.409 M parameters, outperforms recent LSTM-based models and baselines, achieving superior speech quality and intelligibility. Specifically on the WSJ0 dataset, CDCRN delivers a 25.08% improvement in STOI (93.33%) and a 1.52-point increase in PESQ (3.25), with strong results under seen and unseen noise conditions.