<p>The progress of speech recognition systems in high-resource languages exceeds the performance capabilities of low-resource languages, whose data annotations and pre-trained models are insufficient. This research develops VoiceX, an end-to-end recognition model for minimal-resource speech languages, utilising transfer learning mechanisms to address data scarcity issues. Specific domain data is applied to the top of constrained multilingual models to improve speech recognition accuracy in languages with limited resources, as done in VoiceX. Robust feature extraction and sequence modelling in the proposed model rest on its combination of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). VoiceX utilises transfer learning to adapt to limited-resource languages, while still maintaining performance that reaches near-high-resource language standards. Experimental testing indicates that VoiceX manages to lower the Word Error Rate by an average of 33%, while delivering WER performances of 25.4% for Swahili and 22.8% for Nepali, which exceed those of robust baseline systems. VoiceX exhibits enhanced noise tolerance because it provides a 36.6% WER decrease on clear speech and 34.7% on speech obscured by music noise. According to the model evaluations, the Character Error Rate (CER) receives a 32% average reduction. The performance breakthroughs decrease the requirement for substantial labelled datasets, which makes VoiceX operable for digital language protection and promotion of underrepresented languages. The paper evaluates the existing approach before proposing future directions, which include the development of Transformer-based architectures and the expansion of language and dialect capabilities for the model.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VoiceX: An End-to-End Speech Recognition Model for Low-Resource Languages Using Transfer Learning

  • Wided Bouchelligua

摘要

The progress of speech recognition systems in high-resource languages exceeds the performance capabilities of low-resource languages, whose data annotations and pre-trained models are insufficient. This research develops VoiceX, an end-to-end recognition model for minimal-resource speech languages, utilising transfer learning mechanisms to address data scarcity issues. Specific domain data is applied to the top of constrained multilingual models to improve speech recognition accuracy in languages with limited resources, as done in VoiceX. Robust feature extraction and sequence modelling in the proposed model rest on its combination of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). VoiceX utilises transfer learning to adapt to limited-resource languages, while still maintaining performance that reaches near-high-resource language standards. Experimental testing indicates that VoiceX manages to lower the Word Error Rate by an average of 33%, while delivering WER performances of 25.4% for Swahili and 22.8% for Nepali, which exceed those of robust baseline systems. VoiceX exhibits enhanced noise tolerance because it provides a 36.6% WER decrease on clear speech and 34.7% on speech obscured by music noise. According to the model evaluations, the Character Error Rate (CER) receives a 32% average reduction. The performance breakthroughs decrease the requirement for substantial labelled datasets, which makes VoiceX operable for digital language protection and promotion of underrepresented languages. The paper evaluates the existing approach before proposing future directions, which include the development of Transformer-based architectures and the expansion of language and dialect capabilities for the model.