Exploiting Visual Lip Movement Information for Automatic Speaker Recognition
摘要
Automatic speaker recognition (ASR) systems highly rely on the conventional voice biometric modality. Speech generation is fundamentally bimodal as it involves audio as well as visual representations. Lip movement information unlike audio signals is invariant to acoustic noise perturbation. This research highlights the significance of the visual information presented by the unique lip motion in addressing the robust design of an ASR system. This work has employed classical techniques and has compared the outcomes with those derived with application of the state of art neural networks. The reliable Histogram of Gradient feature detector algorithm is employed for extracting lip movement information. Facial Landmarks was further used to extract dynamic lip movements to develop a robust real time system. Traditional Gaussian Mixture models model was chosen for classification of lip movement features and the state-of-the-art Convolution Neural Networks classifier was deployed to compare the results. The contribution of the research conducted is to highlight the significance of dynamic lip movement detection to real time ASR. The system designed attains 91.4% accuracy.