Lipreading, the ability to understand speech by watching a speaker's lip movements, has been a long-standing research area in the field of speech recognition. It is a useful technique for people who are deaf or hard of hearing, as well as for noisy environments where traditional speech recognition systems may not work effectively. The proposed system is a part of deep learning and computer vision fields. In the existing system, the main disadvantage is the model being unable to function on a wide set of vocabulary. The proposed system tackles this problem by having the model train on a large and sophisticated dataset. This project aims to develop a deep learning lipreading model that is capable of mapping a variable-length sequence of video frames to text. The model uses spatiotemporal convolutions to analyse and process spatio-temporal visual features. The model is trained by minimizing connectionist temporal classification loss which measures the difference between predicted and actual outputs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Silent Speech Recognition: Automatic Lip Reading Model Using 3D CNN and GRU

  • T. Mallika Devi,
  • Siripurapu Keerthana,
  • Pentyala Santhi,
  • Puram Pravallika,
  • Sama Rajeshwari

摘要

Lipreading, the ability to understand speech by watching a speaker's lip movements, has been a long-standing research area in the field of speech recognition. It is a useful technique for people who are deaf or hard of hearing, as well as for noisy environments where traditional speech recognition systems may not work effectively. The proposed system is a part of deep learning and computer vision fields. In the existing system, the main disadvantage is the model being unable to function on a wide set of vocabulary. The proposed system tackles this problem by having the model train on a large and sophisticated dataset. This project aims to develop a deep learning lipreading model that is capable of mapping a variable-length sequence of video frames to text. The model uses spatiotemporal convolutions to analyse and process spatio-temporal visual features. The model is trained by minimizing connectionist temporal classification loss which measures the difference between predicted and actual outputs.