Enhanced Malay Automatic Speech Recognition Through Fine-Tuning Pre-trained Whisper Models
摘要
Automatic Speech Recognition (ASR) technology has changed the interaction between humans and machines as evidenced by the creation of voice-controlled devices and real-time transcription. While these improvements are impressive, there are challenges when ASR technology is used in low-resource languages like Malay. The arduous work involved in collecting sufficient labelled Malay speech data and the lengthy model building process have hampered the application of ASR in low-resource languages. A solution to these issues is to perform fine-tuning on a pre-trained model. In this work, a fine-tuning method is proposed to fine-tune Whisper, a recent high-performing ASR model. The proposed fine-tuning process significantly boosts the performance of the ASR model for Malay speech data. In one of the datasets that are used to test our fine-tuned model, the word error rate (WER) has been remarkably reduced from 54.77% to 8.89%.