Speech Enhancement Based on Two-Stage Neural Network with Structured State Space for Sequence Transformation
摘要
In this paper, a new method for improving speech quality using the Structured State Space for Sequence (S4) transformation was proposed. This method inherits existing two-stage denoising methods using recurrent neural networks. However, the use of S4 layers instead of long-term short-term memory brought improvements in two ways. Firstly, it was possible to achieve a reduction in the number of trained parameters of the neural network, while maintaining the quality of speech enhancement. Secondly, due to the use of the convolutional representation of S4 transformations, the network training time per one epoch has decreased. The proposed two-stage neural network model for denoising was implemented using the PyTorch library. For training and testing, a standard DNS Challenge 2020 dataset was used. The optimal type of the loss function for training, and the best number of S4 layers was selected. Comparison with existing real-time speech enhancement methods showed that the developed model was one of the best performers for all quality metrics.