Speech Enhancement Method Based on Fusion Attention with Local Recurrence
摘要
In scenarios with complex background noise, existing speech enhancement models suffer from insufficient denoising performance. To address this challenge, this paper proposes a local recurrent speech enhancement method based on a fusion attention mechanism, aimed at addressing this issue. Building on the classic U-net network architecture, the method incorporates a fusion attention structure within both the encoder and decoder. This structure enables the model to focus on the key information of the overall speech, filtering out noise or interference in specific areas, and enhancing the model's robustness to noise. In the bottleneck layer, a squeeze-and-excitation local recurrent module is designed to improve the model's sensitivity to changes in speech, maintaining the natural fluidity of the speech signal. The quality and intelligibility of the enhanced speech are evaluated on the VoiceBank-DEMAND dataset using five common metrics. Experimental results show that the proposed method outperforms other models across all five speech enhancement metrics.