SAPNN: Self-attention Pyramidal Neural Network for Face Emotion Recognition with Multimodal Fusion
摘要
Emotion detection methods that use many modalities simultaneously have been discovered to be more precise and robust compared to those that depend on a single sense. Because sentiments may be communicated in several ways, each gives a unique and complementary view of the speaker's thoughts and emotions. Thus, combining and analyzing data from several modalities might reveal a person's emotional state. Multimodal fusion-based neural networks integrate voice and physiological information to improve face emotion identification. This strategy uses each modality's capabilities to capture subtle and nuanced data that single-modality systems may overlook, improving the knowledge of human emotions. Here, we adopted a Self-Attention Pyramidal Neural Network (SAPNN) for face emotion recognition with multimodal fusion to enhance the accuracy and robustness of emotion classification. Our SAPNN architecture integrates a pyramidal structure that extracts hierarchical features at multiple scales, capturing both global and local facial details. Within each pyramid level, self-attention mechanisms are employed to focus on critical facial regions and dependencies dynamically, enhancing the network's ability to discern subtle emotional expressions. Experimental results demonstrate that our SAPNN outperforms traditional approaches in terms of accuracy by achieving 98.4% of WAA, 95.2% of WA-F1, 13.53% of MSE, 17.46% of RMSE and 0.98 of PCC.