LipMSTA: Multi-scale Spatio-Temporal Attention for Lip-Reading
摘要
Recent studies on lip-reading typically combine a 3D convolution with ResNet-18 as the front-end network to extract short-term temporal dynamics and spatial features of the lip region. However, these methods are limited in extracting fine-grained spatio-temporal features. For word-level lip-reading task, since lip-reading recognition relies on continuous subtle changes in key region of the video, the network needs to capture more precise spatial location information and subtle dynamic changes in the lip region. In this paper, we introduce the Multi-scale Spatio-Temporal Attention (MSTA) module into the front-end network, which enhances the front-end network’s ability to capture fine-grained spatio-temporal features. In addition, the MSTA module uses the attention mechanism to assign weights to features in different temporal and spatial locations, achieving precise focus on key spatio-temporal regions. Furthermore, we explore the effective integration of the MSTA module into the front-end network. Through a series of experiments, we have determined the optimal structure of the LipMSTA lip-reading model. Our LipMSTA model outperforms all baseline methods and achieves a new state-of-the-art performance on the Lip Reading in the Wild (LRW) dataset.