错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Study of Kazakh Speech Recognition in Hiformer Model

  • Orken Mamyrbayev,
  • Turdybek Kurmetkan,
  • Dina Oralbekova,
  • Nurdaulet Zhumazhan

摘要

This article presents an overview of automatic speech recognition (ASR) technologies and describes the use of an advanced version of the Transformer model, the Hiformer model, in Kazakh speech recognition. A literature review of Kazakh speech recognition systems was made. The structure of the Hiformer model is described and how it can be used in different parts of the structure of an advanced attention mechanism (AED) (encoder, decoder, cross-coder attention). An experiment was carried out on the execution of tasks of recognition of Kazakh speech using the Hiformer model. This study also details an experiment conducted to assess the efficacy of the Hiormer model in recognizing Kazakh speech. The experimental results clearly demonstrate the superiority and enhanced efficiency of the HiFormer model over its predecessors, the Transformer and Conformer models, in handling the complexities of Kazakh speech recognition. This finding underscores the potential of the HiFormer model in advancing ASR technology for languages with unique linguistic characteristics. In recognizing Kazakh speech, the use of the Hiformer model reduced the word error rate (WER) by 3.7% and the character error rate (CER) by 2.2% compared to the Transformer model.