错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Joint Speech and Noise Estimation Using SNR-Adaptive Target Learning for Deep-Learning-Based Speech Enhancement

  • Xiaoran Li,
  • Zilu Guo,
  • Jun Du,
  • Chin-Hui Lee,
  • Yu Gao,
  • Wenbin Zhang

摘要

In this paper, we propose an SNR-adaptive target learning strategy and apply it to a joint speech and noise estimation network to address the mismatch between speech enhancement (SE) and automatic speech recognition (ASR) modules. The progressive learning (PL) methods have revealed the importance of retaining residual noise in the training targets of the enhancement model to alleviate this mismatch. Inspired by this, we adopt an SNR-adaptive target learning strategy to optimize the SNR targets for the SE model, thereby achieving adaptive denoising of the enhancement model in a data-driven manner and further improving its performance on the backend ASR task. Next, we extend the SNR-adaptive target learning strategy to a joint speech and noise estimation network and validate the adaptability of the target learning strategy with the noise prediction branch. We demonstrate the effectiveness of our proposed method on a public benchmark, achieving a significant relative word error rate (WER) reduction of approximately 37% compared to the WER results obtained from unprocessed noisy speech.