SFERNet: Student Facial Expression Recognition Using Superpixel-Assisted Global Semantic Enhancement and Fine-Grained Features
摘要
Facial expression recognition is a challenging task due to large inter-class and small intra-class similarities. Current research still suffers from capturing non-local dependencies with semantic continuity and representing fine-grained local features. To tackle these issues, we propose SFERNet, a novel emotion computation network based on semantic-enhanced global dependency modeling and fine-grained feature compensation. Firstly, we introduce a new superpixel-aware attention module (SPAA) based on superpixel interactions. SPAA aggregates similar semantics through superpixels, reducing facial structure fragmentation and irrelevant information interference in attention computations. This helps the model obtain a global representation efficiently and effectively even at shallow layers. Additionally, to enhance compensation for fine-grained local features, we design the Dual Attention Refinement FFN (DARF) module and Mixture of Expert Feature Extractor (MEFE). DARF refines local features through fine-grained interactions across spatial positions and channels, dynamically fusing features for high-level image representation. MEFE, with multiple experts working in parallel, facilitates adaptive capturing of multi-scale local features and low-level details. Moreover, to alleviate the scarcity of student facial images in classroom settings, we introduce a new student facial expression dataset, SFERD. Ablation and comparison experiments demonstrate SFERNet's superiority over other state-of-the-art methods across multiple datasets.