Triple-Path RNN Network: A Time-and-Frequency Joint Domain Speech Separation Model
摘要
Studies in speech separation have achieved significant success in recent years. To correctly separate the mixture signals, it is critical to encode the signals into an appropriate latent space. Existing speech separation methods include transforming mixed signals into frequency domain space or time domain space. The frequency domain features (spectrogram) are generated by STFT, which is closely related to speech articulation and reflects the energy of speech directly. The time domain features are learned from a latent embedding space, and the separation effect is facilitated by the end-to-end structure. However, these methods are based on the representations from only one domain, which is insufficient for providing a speech separation encoding space that is completely separable. Therefore, a Triple-Path Recurrent Neural Network (TPRNN) that fuse features from two domains is proposed. It employs a spectrogram as auxiliary information to improve the performance of speech separation. Experimental results on the Wall Street Journal (WSJ0) dataset show that this approach is beneficial to improve speech separation performance.