Hidden Markov Model Based Text to Speech Synthesis for Afan Oromo
摘要
The purpose of text-to-speech synthesis systems is to generate understandable and natural-sounding speech. For many years, this has been a challenging task. Various strategies have been researched and tried to get around these obstacles. Although speech synthesizers based on the hidden Markov model are used for languages other than Afan Oromo, these synthesizers do not take the language's unique characteristics into account. As a result, this research employs Hidden Markov Model-based text-to-speech synthesis for Afan Oromo. An amount of 527 sentences are used for setting up the model from a corpus having a size of 10,112 sentences. The system's performance is evaluated using a total of 20 sentences that are not part of the training dataset. The Mean opinion score evaluation method is utilized in this study. The mean opinion score for intelligibility was 4.04, and the mean opinion score for naturalness was 3.51. The synthesis system is rated fair in terms of naturalness and good in terms of intelligibility based on the mean opinion score results. The outcome acquired is empowering and future research bearings are proposed to work on the exhibition of the framework.