Deep hybrid architectures and DenseNet35 in speaker-dependent visual speech recognition
摘要
Visual speech recognition (VSR) translates the visual speech cues into transcription. Speaker-dependent VSR (SD-VSR) can be used for authentication and secure human-computer interactions, where the system only has to recognize a legitimate user's visual speech. This paper presents hybrid deep learning architectures and DenseNet35 for SD-VSR. Two main objectives guide this study to improve SD-VSR accuracy: (1) Designing end-to-end trainable ResNet18-based deep hybrid architectures to investigate suitable input types among optical flow, XCS-LBP, contour, depth, and newly proposed lip-signature along with the RGB. Therefore, 2D and 3D-ResNet18 were modified and used to build hybrid architectures with a late fusion network consisting of densely connected networks and a weighted addition layer, whose weights were learnt during training. (2) Designing customized deep neural network front-end architecture with fewer network parameters and analyzing the impact of proposed video augmentations and dim-light enhancement techniques applied for the training set to improve system's accuracy. Therefore, novel lightweight architectures, 2D and 3D-DenseNet35, are designed with fewer network parameters than DenseNet201 and DenseNet121, reducing the architecture's computational complexity while maintaining the performance. The MIRACL-VC1 benchmark was used to assess the performance of all designed networks, as the dataset included both RGB and depth videos. Experimental results show that the proposed 2D and 3D-DenseNet35 with video augmentations surpasses all other SD-VSR frameworks and achieves state-of-the-art performance by obtaining 98% average test accuracy over five folds. Thus, the proposed video augmentations successfully address the challenges associated with VSR and improve its efficiency by preventing overfitting.