Design and Development of a Working Tool for Visual Speech and Speaker Recognition for Marathi and Gujarati Languages
摘要
Multimodal speech recognition is a rapidly growing field of research that aims to combine visual and auditory information to improve the accuracy of speech recognition. This is motivated by the fact that humans use both visual and auditory cues to understand speech and that combining these cues can help to overcome the limitations of either modality alone. In this paper, we present a new approach to multimodal speech recognition that focuses on the visual cues provided by the lips. The lips are a rich source of information about speech, and they can be used to track the movements of the mouth and to identify the phonemes being spoken. We propose a new method for tracking the lip boundary and lip area, and we evaluate the performance of our approach on a standard dataset of lip videos. We also investigate the use of different classification approaches for multimodal speech recognition, including principle component analysis, support vector machines, hidden Markov models, deep learning, and the Viterbi algorithm. Our results show that our approach to lip tracking is effective, and that multimodal speech recognition can achieve significant improvements over traditional audio-only speech recognition.