Dynamic Occlusion Expression Recognition Based on Improved GAN
摘要
In order to address the issue of local occlusion in practical dynamic expression recognition, this paper first introduces a facial restoration network that combines Vision Transformer (ViT) and GAN. This network can accurately identify missing facial features and perform detailed and efficient restoration. Secondly, for the task of expression recognition, a more robust dynamic expression recognition network is trained by cascading ViT with a Two-Stream CNN, effectively leveraging ViT’s feature extraction capability and the Two-Stream CNN’s ability to acquire spatio-temporal features. Finally, by combining these two networks, we can efficiently recognize dynamically occluded expressions. A multitude of experiments demonstrate that the facial image restoration network trained on the CelebA and VGG Face2 datasets outperforms other networks in handling small and medium occlusions. Expression recognition experiments on AFEW and MMI datasets show that this paper’s expression recognition network achieves an accuracy of 54.95% and 81.2%, respectively, for dynamic expression recognition, surpassing mainstream networks. Moreover, the restoration network outperforms mainstream networks in addressing occlusions and provides an average accuracy improvement of 5.34% in occluded expression recognition.