Toward Developing Attention-Based End-To-End Automatic Speech Recognition
摘要
In recent decades, significant research has been conducted in the field of Automatic Speech Recognition (ASR). Machine learning in the field of Natural Language Processing (NLP) is an innovative area of research and has become a popular topic for researchers. This paper aims to examine various Automatic Speech Recognition systems developed over the past decades. Antiquated ASR systems consist of acoustic model, language model, pronunciation model, Weighted Finite-State Transducer (WFST)-based decoder, and text-normalizer. These components are separately trained and assembled. We will discuss various end-to-end (E2E) structures such as Connectionist Temporal Classification (CTC), Recurrent Neural Network-Transducer (RNN-T), and the attention-based model. We will also explore some of the modern attention-based architectures such as Jasper, Transformer, and Conformer model.