错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Speech Recognition for Telugu Language Using Pre-trained Wav2Vec 2.0 Model

  • Sri Gani Kaarthikeya Kammula,
  • Gunasekhar Devineni,
  • Chiranjeevi Karanki,
  • Muzaffar Ahmad Dar,
  • Jagalingam Pushparaj

摘要

This study focuses on creating a system for automatic recognition of speech tailored for Telugu, a spoken language in India. Deep learning (DL) techniques have been extensively utilized to create ASR systems across diverse languages and domains in recent years. These models, however, require substantial training resources, particularly large corpora of continuous speech utterances collected from multiple speakers, along with their corresponding transcripts. Unfortunately, such comprehensive speech datasets are often unavailable for many Indian languages, including Telugu. This study utilizes a pre-trained Wav2Vec 2.0 model for developing an ASR system in the Telugu language, addressing the challenge of limited data availability and demonstrating remarkable results even when fine-tuned on a limited dataset. We fine-tuned Wav2Vec 2.0 on OpenSLR and Mozilla Common Voice, achieving a 25.6% word error rate (WER) and a 4.5% character error rate (CER) on OpenSLR, and an 18% CER on Mozilla Common Voice. These results demonstrate its efficacy in languages with data disparities and highlight the effectiveness of leveraging pre-trained architectures in low-resource linguistic settings.