Lexical Interpretation of Visual Cues Using Deep Learning
摘要
Lexical interpretation of visual cues is an approach for understanding spoken phrases by visually observing the movements and shapes of a speaker’s lips. A comprehensive review of the existing methods exposes the limitations of traditional lip reading techniques in capturing both spatial and temporal dimensions of lip movements. To address this gap, this project presents an approach to advance lip reading efficacy by synergizing Convolutional Neural Networks (CNN) and Gated Recurrent Units (GRU). The system has achieved an accuracy of 97% on GRID dataset. The implications of this research extend to improved communication accessibility for individuals with hearing impairments as well as broader applications in areas such as criminal investigations and security.