Micro-Expression Recognition via CNN and Multi-path Vision Transformer Integrated with Spatiotemporal Separated Self-attention
摘要
Deep learning-based methods have exhibited outstanding performance in micro-expression recognition. However, convolutional operations are performed on fixed-size windows, which limits their ability to learn long-range relationships between different facial action units. To address this issue, we propose a novel Micro-Expression Recognition via CNN and Multi-Path Vision Transformer Integrated with Spatiotemporal Separated Self-Attention (CMVT). Among them, the combination of CNN and Vision Transformer is adopted to learn the local features and long-distance dependency of micro-expression facial action units. In addition, we use the motion magnification algorithm to amplify the motion amplitude of micro-expression apex frame for extracting more discriminative features. Experimental results on various datasets verify the effectiveness of the presented approach compared to the existing micro-expression methods.