A Subconvolutional U-net with Gated Recurrent Unit and Efficient Channel Attention Mechanism for Real-Time Speech Enhancement
摘要
We propose a subconvolutional U-net with a gated recurrent unit and an efficient channel attention mechanism for real-time speech enhancement. The subconvolutional U-net (SCU-net) is a convolutional encoder−decoder model with skip connections. The SCU-net encoder block is a combination of the multi-scale convolutional block (MCB), and the feature calibration (FC) block. MCB uses subconvolutions with different sizes of kernels to extract global and local contextual feature information. In FC, calibration coefficients are extracted by determining the nonlinear relationship within the multi-dimensional features and assigning a relatively higher weight to speech components and a lower weight to noise components within the feature. Efficient channel attention (ECA) can implement a cross-channel interaction without reducing its dimensions. Network performance was significantly improved by using a customized kernel size in module tests. The SCU-net decoder block is a replica of the encoder block. Additionally, SCU-net uses gated recurrent units (GRU), for learning long-range dependencies. It is noise and speaker-independent, so noise types and speakers can differ between training and testing. Compared to LSTM-based models, SCU-GRU-ECANet consistently leads to better objective comprehensibility and perceptual quality.