<p>Automatic speaker recognition (ASR) systems highly rely on the conventional voice biometric modality. Speech generation is fundamentally bimodal as it involves audio as well as visual representations. Lip movement information unlike audio signals is invariant to acoustic noise perturbation. This research highlights the significance of the visual information presented by the unique lip motion in addressing the robust design of an ASR system. This work has employed classical techniques and has compared the outcomes with those derived with application of the state of art neural networks. The reliable Histogram of Gradient feature detector algorithm is employed for extracting lip movement information. Facial Landmarks was further used to extract dynamic lip movements to develop a robust real time system. Traditional Gaussian Mixture models model was chosen for classification of lip movement features and the state-of-the-art Convolution Neural Networks classifier was deployed to compare the results. The contribution of the research conducted is to highlight the significance of dynamic lip movement detection to real time ASR. The system designed attains 91.4% accuracy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploiting Visual Lip Movement Information for Automatic Speaker Recognition

  • Sumita Nainan,
  • Sonal Parmar,
  • Kanchan Bakade

摘要

Automatic speaker recognition (ASR) systems highly rely on the conventional voice biometric modality. Speech generation is fundamentally bimodal as it involves audio as well as visual representations. Lip movement information unlike audio signals is invariant to acoustic noise perturbation. This research highlights the significance of the visual information presented by the unique lip motion in addressing the robust design of an ASR system. This work has employed classical techniques and has compared the outcomes with those derived with application of the state of art neural networks. The reliable Histogram of Gradient feature detector algorithm is employed for extracting lip movement information. Facial Landmarks was further used to extract dynamic lip movements to develop a robust real time system. Traditional Gaussian Mixture models model was chosen for classification of lip movement features and the state-of-the-art Convolution Neural Networks classifier was deployed to compare the results. The contribution of the research conducted is to highlight the significance of dynamic lip movement detection to real time ASR. The system designed attains 91.4% accuracy.