Leveraging 3D-CNN and graph neural network with attention mechanism for visual speech recognition
摘要
Deep learning techniques have demonstrated early advancements in addressing the challenges of complex Visual Speech Recognition (VSR) tasks. Nonetheless, a persistent issue arises when distinguishing characters or words with similar pronunciations, known as homophones, which results in ambiguity. Existing VSR systems also face technical constraints due to insufficient visual data for learning short-duration phonemes like “at”, “an”, “a”, and “eight”. Moreover, cutting-edge VSR techniques perform exceptionally well when interpreting overlapping speakers. However, extending these methods to unseen speakers leads to a significant performance decline due to the limited diversity in the training dataset and substantial variations in physical attributes, such as lip shape and color, across different speakers. To address the existing challenges in VSR, we propose a multi-modal approach that leverages visual and landmark information to capture complex spatio-temporal patterns for the model generalization capabilities. The model employs a multi-layered Three-Dimensional Convolutional Neural Network (3D-CNN) that extracts visual features, while a Graph Convolutional Network (GCN) captures precise landmark information for accurate lip shape localization. The extracted features are then fused for further processing using a Sequence-to-Sequence (Seq2Seq) model based on the attention mechanism. The proposed model achieved a WER of 0.53% and 8.21% for the overlap and unseen speakers category. Notably, these results surpass the performance of existing models, demonstrating remarkable accuracy for VSR on the GRID dataset in both the unseen and overlapping speaker scenarios.