Advancing Lip-Reading for Unseen Speakers Through Fusion and Augmentation of Spatio-temporal Landmarks and Visual Features
摘要
Lip reading decodes spoken language by observing lip movements, improving speech recognition in noisy environments. Current techniques rely primarily on visual features, which often result in suboptimal generalization to unseen speakers. Lip landmark features, which describe speaker-independent lip movements, can overcome this limitation. This paper presents a novel approach to constructing spatio-temporal features from landmarks. By using intra-frame Euclidean distances along with inter-frame velocities and accelerations, it effectively captures lip movements, thereby improving the generalization of lip-reading models to unseen speakers. The proposed multimodal fusion architecture integrates visual and landmark features through cross-attention mechanisms, following self-attention processing in each branch. Residual connections integrate pre- and post-fusion features to enhance comprehensive feature representation, and a feature enhancement strategy using random feature masking further augments the model’s generalization capabilities. A new dataset, CECC-VSR, is constructed, comprising a limited number of speakers recorded in real, complex indoor or outdoor environments, where speakers can freely move within the recording scenes. Experimental results on the GRID, LRW-ID, and CECC-VSR datasets indicate that the proposed method significantly improves generalization for unseen speakers. It reduces the WER by 4.16% on the GRID and increases accuracy by 1.97% on the LRW-ID, both compared to the baseline.