Joint Time-Domain and Frequency-Domain Progressive Learning for Single-Channel Speech Enhancement and Recognition
摘要
Single-channel speech enhancement for automatic speech recognition (ASR) has been extensively researched. Traditional methods usually directly learn clean target, which may introduce speech distortions and limit ASR performance. Meanwhile, these methods usually focus on either the time or frequency domain, ignoring their potential connections. To tackle these problems, we propose a joint time and frequency domain progressive learning (TFDPL) method for speech enhancement and recognition. TFDPL leverages information from both domains to estimate frequency masks and waveforms, and further combines the information from both domains through a fusion loss, gradually predicting less-noisy and cleaner targets. Experimental results show that TFDPL outperforms traditional methods in ASR and perceptual metrics. TFDPL achieves relative reductions of 43.83% and 36.03% in word error rate for its intermediate outputs on the CHiME-4 real test set using two different acoustic models and certain improvements in PESQ and STOI metrics for clean output on the simulated test set.