Frequency-prior enhanced network for facial expression recognition via dynamic large kernels and dual-domain learning
摘要
Facial Expression Recognition in natural scenes constitutes a critical research direction in affective computing. Prevailing approaches predominantly rely on RGB spatial domain modeling, neglecting the inherent texture prior information in the frequency domain, which consequently limits their adaptability to challenging scenarios. Furthermore, conventional methods typically depend on single-domain spatial features, failing to effectively exploit the complementary characteristics between spatial and frequency domains, thereby limiting the model’s capacity for subtle expression representation. To address these limitations, this paper proposes FPNet, a dual-domain collaborative framework for FER. Specifically, we design three core components: (1) DyLKBlock constructs dynamic spatial feature extraction through cascaded large-kernel convolutions, achieving an equivalent 23