A Study of Kazakh Speech Recognition in Hiformer Model
摘要
This article presents an overview of automatic speech recognition (ASR) technologies and describes the use of an advanced version of the Transformer model, the Hiformer model, in Kazakh speech recognition. A literature review of Kazakh speech recognition systems was made. The structure of the Hiformer model is described and how it can be used in different parts of the structure of an advanced attention mechanism (AED) (encoder, decoder, cross-coder attention). An experiment was carried out on the execution of tasks of recognition of Kazakh speech using the Hiformer model. This study also details an experiment conducted to assess the efficacy of the Hiormer model in recognizing Kazakh speech. The experimental results clearly demonstrate the superiority and enhanced efficiency of the HiFormer model over its predecessors, the Transformer and Conformer models, in handling the complexities of Kazakh speech recognition. This finding underscores the potential of the HiFormer model in advancing ASR technology for languages with unique linguistic characteristics. In recognizing Kazakh speech, the use of the Hiformer model reduced the word error rate (WER) by 3.7% and the character error rate (CER) by 2.2% compared to the Transformer model.