Task-Adaptive Generative Adversarial Network Based Speech Dereverberation for Robust Speech Recognition
摘要
Reverberation is known to severely affect speech recognition performance when speech is recorded in an enclosed space. Deep learning-based speech dereverberation has been remarkably successful in recent years, achieving superior recognition performance for far-field speech applications. However, the output from conventional dereverberation systems cannot be guaranteed suitable for back-end recognition systems because of their different task goals. To bridge the gap between the front-end dereverberation and the back-end recognition, we propose a novel task-adaptive speech dereverberation generative adversarial network (GAN) based speech dereverberation model called Task-adaptive GAN. Specifically, we propose to replace the binary-valued discriminator in a regular generative adversarial network with a novel senone-predicted discriminator, and also introduce a well-designed recognition-aware generator as a dereverberation system. By doing so, the corresponding output distribution will be more suitable for the recognition task. Experimental results on the REVERB corpus show that our proposed approach achieves a relative 18.6% and 8.6% word error rate reduction than the traditional GAN-based baseline system on the simulated set and real set, respectively.