With the popularization of monitoring devices, video face recognition has gradually become a new research direction. However, single image based excellent deep recognition models often face the challenge of performance degradation when dealing with video face recognition task directly. Moreover, these deep models often need huge amounts of data for training, while the publicly available video face dataset is relatively limited, resulting in poor model robustness and weak generalization ability, seriously affecting the practical application of the model. In response to this issue, a new deep video face recognition framework via integrating YoLov4 and mask autoencoder (MAE) networks (IYMN) is proposed in this paper, which has stronger robustness on limited video datasets. Specifically, the proposed IYMN first uses YoLov4 for high-precision and fast face detection, then uses MAE to learn robust frame-level features, finally uses the Time Transformer network to obtain the video-level features. The experimental results show that the proposed method achieves satisfactory performance on public video face datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

IYMN: A New Video Face Recognition Method Via Integrating YoLov4 and MAE Networks

  • Feng Zhou,
  • Xiangfang Ji

摘要

With the popularization of monitoring devices, video face recognition has gradually become a new research direction. However, single image based excellent deep recognition models often face the challenge of performance degradation when dealing with video face recognition task directly. Moreover, these deep models often need huge amounts of data for training, while the publicly available video face dataset is relatively limited, resulting in poor model robustness and weak generalization ability, seriously affecting the practical application of the model. In response to this issue, a new deep video face recognition framework via integrating YoLov4 and mask autoencoder (MAE) networks (IYMN) is proposed in this paper, which has stronger robustness on limited video datasets. Specifically, the proposed IYMN first uses YoLov4 for high-precision and fast face detection, then uses MAE to learn robust frame-level features, finally uses the Time Transformer network to obtain the video-level features. The experimental results show that the proposed method achieves satisfactory performance on public video face datasets.