A Fusion Approach for Kannada Speech Recognition Using Audio and Visual Cue
摘要
The act of conveying thoughts and ideas to another person without using words or visible facial expressions is known as communication. It is a vital component of human existence since it enables interaction with the outside world. An audiovisual speech recognition system that uses image processing to aid speech recognition systems in the lip-reading process is proposed in this work. This helps the deaf and hearing-impaired by identifying the word using the Kannada audio database, a Dravidian language spoken by more than 60 million people in the state of Karnataka. Three components make up the proposed speech recognition system: the audio model, the visual model, and the fusion model. Mel-frequency Cepstrum coefficient is successfully used to extract audio information, and random forest employed a classification and achieved an accuracy of 86%. The feed-forward neural network was employed as a classifier and achieved a 57% accuracy rate. The Viola-Jones technique was utilised to extract the visual features. With the help of an artificial neural network and random forest algorithm, audio and visual fusion achieved an accuracy of 89%. It demonstrates that when the audio and the visual models are combined, the accuracy is higher than when each model is used alone.