Comparative Analysis of Models for Neural Machine Speech-to-Text Translation for Turkic State Languages
摘要
In this work, we compare and evaluate speech recognition models for the Turkic state languages, namely Azerbaijani, Kazakh, Kyrgyz, Turkish, Turkmen, and Uzbek. For this purpose, experimental studies of neural speech recognition are being conducted for three available open-source models: Whisper is an ASR system by OpenAI, TurkicASR of ISSAI, and The Massively Multilingual Speech (MMS) project of Facebook AI’s initiative. This project represents a key step towards streamlining the process of recording and processing meeting minutes in diverse Turkic languages. The scientific contribution of this article is the comparative analysis and selection of speech recognition models for the Turkic state languages based on ongoing experimental studies.