错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging Wav2Vec2.0 for Kazakh Speech Recognition: An Experimental Study

  • Zhanibek Kozhirbayev

摘要

In the fast-growing world of neural networks, models trained on extensive multilingual text and speech data have shown great promise for improving the state of low-resource languages. This study focuses on the application of state-of-the-art speech recognition models, specifically Facebook’s Wav2Vec2.0 and Wav2Vec2-XLSR, to the Kazakh language. The primary objective is to evaluate the performance of these models in transcribing spoken Kazakh content. Additionally, the research explores the possibility of using data from other languages for initial training and examines whether fine-tuning the model with target language data can improve its performance. More so, this work gives insights into how effective pre-trained multilingual models are when used on low-resource languages. The fine-tuned wav2vec2.0-XLSR model demonstrated impressive results, achieving a character error rate (CER) of 1.9 and a word error rate (WER) of 8.9 when tested against the test set of the Kazcorpus dataset. These findings may help create robustness in Automatic Speech Recognition (ASR) systems for Kazakh which could be used for various applications such as voice-activated assistants; speech-to-text translators among others.