错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An amalgamation of integrated features with DeepSpeech2 architecture and improved spell corrector for improving Gujarati language ASR system

  • Mohit Dua,
  • Bhavesh Bhagat,
  • Shelza Dua

摘要

Automatic Speech Recognition systems that convert language into written text have greatly transformed human–machine interaction. Although these systems have achieved results, in languages building accurate and reliable ASR models for low resource languages like Gujarati comes with significant challenges. Gujarati lacks data and linguistic resources, making developing high-performance ASR systems quite difficult. In this paper, we propose an approach to enhance the effectiveness of a Gujarati ASR model despite resources. We achieve this by incorporating integrated features such as Mel Frequency Cepstral Coefficients (MFCC) and Gammatone Frequency Cepstral Coefficients (GFCC) utilizing the DeepSpeech2 architecture and implementing an improved spell correction technique based on the Bidirectional Encoder Representations from Transformers (BERT) algorithm. Our approach has demonstrated superiority over previous state-of-the-art methodologies through testing and evaluation. The experimental results demonstrate that our proposed method consistently reduces the Word Error Rate (WER) by 10–12 percentage points compared to the existing work, surpassing the most significant improvement of 5.87%. Our findings demonstrate the viability of developing accurate and dependable ASR systems for languages with limited resources, such as Gujarati.