Review of Automatic Speech Recognition Systems for Ukrainian and English Language
摘要
Automatic speech recognition systems are highly regarded today since they can improve inclusivity, streamline business communications, etc. This page overviews the most recent scientific research and a quick description of the automatic speech recognition system’s overall structure. The paper evaluates popular speech recognition products in English and Ukrainian based on model parameters, error frequency, and processing time. Studying information dynamics during internet warfare has led to developing technologies that enhance human–machine interaction, including speech recognition. Speech recognition systems convert spoken language into readable text, enabling convenient and fast communication in various business applications. The accuracy and capabilities of speech recognition systems depend on the underlying algorithms and technologies employed. This paper provides an overview of the leading technologies in speech recognition, including hidden Markov models, natural language processing, N-grams, and artificial intelligence. The traditional hybrid approach to automatic speech recognition (ASR) has been widely used, but it has limitations regarding accuracy and training time. End-to-end deep learning has emerged as a promising alternative for speech recognition, directly converting audio input to text. This paper overviews commercial and open-source ASR tools, highlighting popular options like DeepSpeech, Whisper, and Facebook Wav2Vec 2.0. Open-source ASR systems offer flexibility and integration possibilities, making them an attractive option for developers. This paper offers insights into the advancements and challenges in automatic speech recognition, which covers various technologies, research studies, and software tools that contribute to improving the accuracy and functionality of speech recognition systems.