Emotion-aware adaptation of CLIP model for facial expression recognition
摘要
Facial expression recognition (FER) remains a challenging task due to subtle variations in facial details and unconstrained conditions such as changes in head posture, illumination, and occlusion. Current FER approaches primarily focus on capturing discriminative facial features in vision manner, often neglecting the rich semantic information available in textual modalities. Additionally, these methods typically rely on generic classification templates, which fail to capture instance-specific features, resulting in inadequate representation and fine-grained discrimination ability. To tackle the above issues, we propose a novel emotion-aware adaptation framework that integrates the pre-trained CLIP model for FER, leveraging both visual and textual modalities to enhance representation learning and capture fine-grained emotional details. Specifically, we introduce the Expression-aware adapter module to capture emotion-specific facial representations through task-specific fine-tuning while preserving the generalization capabilities of the CLIP model. Furthermore, the instance-enhanced expression classifier module is proposed to enhance textual descriptors with instance-specific visual embeddings using spherical linear interpolation, creating a more precise and discriminative classifier. Extensive experiments on three in-the-wild FER benchmarks demonstrate superiority of our proposed approach.