错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Noise Robustness of Automatic Speech Recognition Based on a Parallel Adapter Model with Near-Identity Initialization

  • Takahiro Osaki,
  • Yui Sudo,
  • Katsutoshi Itoyama,
  • Kenji Nishida,
  • Kazuhiro Nakadai

摘要

This paper proposes the parallel adapter model (PAM) to improve the noise-robustness of automatic speech recognition (ASR) systems with a small amount of retraining. The performance of ASR degrades in noisy environments; thus, speech enhancement (SE) methods are frequently employed as a front-end processor. However, the entire model, including the SE, and ASR, should be retrained to relax the mismatch problem between the SE and ASR, which is a time-consuming process. To solve this problem, the proposed PAM utilizes two small networks (called adapters) and a corresponding adapter initialization technique. The adapters are positioned in the models where mismatch tends to occur frequently. Based on the proposed method, we have implemented ASR models, which was evaluated, and compared experimentally. The experimental results demonstrate that the PAM improved the ASR accuracy with only 7–8 retraining epochs, and the improvement reached 26.1 points with the best performance.