IYMN: A New Video Face Recognition Method Via Integrating YoLov4 and MAE Networks
摘要
With the popularization of monitoring devices, video face recognition has gradually become a new research direction. However, single image based excellent deep recognition models often face the challenge of performance degradation when dealing with video face recognition task directly. Moreover, these deep models often need huge amounts of data for training, while the publicly available video face dataset is relatively limited, resulting in poor model robustness and weak generalization ability, seriously affecting the practical application of the model. In response to this issue, a new deep video face recognition framework via integrating YoLov4 and mask autoencoder (MAE) networks (IYMN) is proposed in this paper, which has stronger robustness on limited video datasets. Specifically, the proposed IYMN first uses YoLov4 for high-precision and fast face detection, then uses MAE to learn robust frame-level features, finally uses the Time Transformer network to obtain the video-level features. The experimental results show that the proposed method achieves satisfactory performance on public video face datasets.