Facial expression recognition (FER) in uncontrolled real-world settings confronts multifaceted challenges such as dynamic lighting changes, partial occlusions, and varying head poses, necessitating improved model robustness and cross-scenario adaptability. While existing Transformer-based FER methods effectively model long-range dependencies in facial expressions, their computational efficiency is severely hindered by the quadratic complexity of self-attention mechanisms. In light of these constraints, we propose a novel Mamba-Based Dual-Perception Network (MambaFER). First, we designed a lightweight dual-branch Global-Local Feature Perception (GLFP) module that extracts global long-range dependency features through SSM and captures local detail features using CNN, while employing channel splitting and shuffling strategies to achieve efficient feature interaction, significantly reducing parameter count and computational costs while enhancing robust representation capabilities for complex expressions. Furthermore, we developed a Facial Salient Feature Perception (FSFP) module combines landmark detection with attention mechanisms to prioritize expression-critical regions, ensuring holistic feature representation. Experiments demonstrate that our approach maintains lightweight efficiency (27.05M Params, 3.62G FLOPs) while surpassing advanced methods on several benchmarks, achieving 91.07% on RAF-DB, 66.05% on AffectNet-7, 62.78% on AffectNet-8, and 90.58% on FERPlus.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MambaFER: A Mamba-Based Dual-Perception Network for Facial Expression Recognition in the Wild

  • Chao Zhang,
  • XueChuan Huang,
  • Ming Fang

摘要

Facial expression recognition (FER) in uncontrolled real-world settings confronts multifaceted challenges such as dynamic lighting changes, partial occlusions, and varying head poses, necessitating improved model robustness and cross-scenario adaptability. While existing Transformer-based FER methods effectively model long-range dependencies in facial expressions, their computational efficiency is severely hindered by the quadratic complexity of self-attention mechanisms. In light of these constraints, we propose a novel Mamba-Based Dual-Perception Network (MambaFER). First, we designed a lightweight dual-branch Global-Local Feature Perception (GLFP) module that extracts global long-range dependency features through SSM and captures local detail features using CNN, while employing channel splitting and shuffling strategies to achieve efficient feature interaction, significantly reducing parameter count and computational costs while enhancing robust representation capabilities for complex expressions. Furthermore, we developed a Facial Salient Feature Perception (FSFP) module combines landmark detection with attention mechanisms to prioritize expression-critical regions, ensuring holistic feature representation. Experiments demonstrate that our approach maintains lightweight efficiency (27.05M Params, 3.62G FLOPs) while surpassing advanced methods on several benchmarks, achieving 91.07% on RAF-DB, 66.05% on AffectNet-7, 62.78% on AffectNet-8, and 90.58% on FERPlus.