Weighted Spatiotemporal Feature and Multi-task Learning for Masked Facial Expression Recognition
摘要
Human facial expressions are a vital form of non-verbal communication that convey significant emotional information. Occasionally, individuals exhibit facial expressions that do not correspond to their genuine emotions, which are termed masked facial expressions (MFEs). Automatic recognition of MFEs can reveal individuals’ real feelings. However, the complexity inherent in MFEs, resulting from various factors, poses a significant challenge for extracting discriminative representations from MFE videos. In this study, we innovatively designed a multi-task learning network to decompose the challenging task of mixed expression recognition task into two simpler sub-tasks, thereby alleviating the learning burden on the network. Furthermore, we propose a method called spatiotemporal feature modulation (STFM), which comprises a spatiotemporal feature extractor (STFE) and feature weighting module (FWM), aiming to enable the network to focus on features that are beneficial for classification. Additionally, we developed an adaptive spatial attention module (ASAM) to eliminate redundant data in videos and enhance model efficiency by leveraging the action information embedded in dynamic images. The experimental results demonstrate that our method excels in most challenging 36-class task, achieving an accuracy improvement of 4.75% and 0.77% on the 36-class task and 6E task respectively compared to the current state-of-the-art methods.