Leveraging 3D Facial Reconstruction to Enhance Micro-Expression Recognition
摘要
Micro-expression recognition is challenging due to limited data and the subtle, short nature of expressions. Traditional 2D methods struggle with lighting variations and lack geometric awareness. In this study, we conduct a benchmark based on common groups of models (Convolutional Neural Network, Vision Transformer and Image Foundation Models) on two widely used micro-expression datasets, CASME II and SAMM. Notably, Image Foundation Models, though trained on massive datasets, still produce suboptimal results, with evaluation metric F1-scores even lower than those of other models such as CNNs and ViTs trained with much less data. In addition, to overcome the limitations of purely appearance-based image cues, we leverage geometry-aware, dynamic 3D facial features derived from monocular 3D reconstruction models. Experimental results show that incorporating 3D information significantly improves the performance of most baseline models, with F1 gains ranging from 1 to 31%. These results confirm the effectiveness of incorporating 3D geometric information for micro-expression recognition and suggest that our proposed 3D-based approach opens up promising directions for future research.