A Mix Fusion Spatial-Temporal Network for Facial Expression Recognition
摘要
Facial expression is a powerful, natural and universal signal for human beings to convey their emotional states and intentions. In this paper, we propose a new spatial-temporal facial expression recognition network which outperforms many state-of-the-art methods. Our model is composed by two networks, a temporal feature extraction network based on facial landmarks and a spatial feature extraction network based on densely connected network. Image preprocessing method is optimized according to the features of the expression image to reduce network’s overfitting on small datasets. In addition, we propose a mix fusion strategy to better combine spatial and temporal features. Finally, experiments on public datasets are conducted to verify the effectiveness of each module and the improvement of expression recognition accuracy of the spatial-temporal fusion network. The accuracies on OULU-CASIA and CK + datasets reach 90.21% and 99.82% respectively.