In the realm of traditional Mandarin automatic speech recognition (ASR), the estimation of transition probabilities holds a pivotal role within the hybrid Gaussian mixture model (GMM)-hidden Markov model (HMM) or deep neural network (DNN)-HMM architectures. Conversely, the utilization of transition probabilities in end-to-end (E2E) approaches for Mandarin speech recognition remains uncommon. Here, we introduce the integration of transition probabilities into E2E Mandarin speech recognition. Our approach commences with an examination of the attention rescoring decoding mode within the WeNet framework. Subsequently, we leverage the discounting method, rooted in add-one smoothing, to compute the transition probabilities for various levels of modeling units. Finally, we incorporate these transition probabilities into the n-best rescoring process within the attention-based encoder-decoder (AED) decoder component. Through experiments conducted on the AISHELL-1 dataset, we demonstrate the superiority of our approach over the baseline conformer-based system, irrespective of the recording conditions being clean or noisy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attention Rescoring with Transition Probabilities for End-to-End Mandarin Speech Recognition

  • YingWei Tan

摘要

In the realm of traditional Mandarin automatic speech recognition (ASR), the estimation of transition probabilities holds a pivotal role within the hybrid Gaussian mixture model (GMM)-hidden Markov model (HMM) or deep neural network (DNN)-HMM architectures. Conversely, the utilization of transition probabilities in end-to-end (E2E) approaches for Mandarin speech recognition remains uncommon. Here, we introduce the integration of transition probabilities into E2E Mandarin speech recognition. Our approach commences with an examination of the attention rescoring decoding mode within the WeNet framework. Subsequently, we leverage the discounting method, rooted in add-one smoothing, to compute the transition probabilities for various levels of modeling units. Finally, we incorporate these transition probabilities into the n-best rescoring process within the attention-based encoder-decoder (AED) decoder component. Through experiments conducted on the AISHELL-1 dataset, we demonstrate the superiority of our approach over the baseline conformer-based system, irrespective of the recording conditions being clean or noisy.