SESNet: A Speech Enhancement and Separation Network in Noisy Reverberant Environments
摘要
Speech enhancement and separation in noisy reverberant environments are very challenging tasks. In this paper, we propose a speech enhancement and separation network, SESNet, for speech enhancement or speech separation in noisy reverberant environments, which is a multi-scale encoder-decoder architecture including a global-local feature extractor (GLFE). We also explored four kinds of Former blocks to be equipped in GLFE. We evaluate the performance of speech enhancement and speech separation on the VoiceBank+DEMAND and the WHAMR! datasets. The experimental results show that the SESNet has excellent performance for single- and multi-channel speech enhancement, and single-channel multi-speaker speech separation, keeping with a small model size.