<p>Micro-expression recognition (MER) remains a challenging task due to the subtle and transient nature of micro-expressions, which often results in poor generalization and low recognition accuracy in existing approaches. This research introduces a novel MER framework that integrates graph convolutional networks (GCNs) with facial action unit features to address critical limitations, including ineffective feature extraction, class imbalance, and insufficient temporal modeling. The proposed method effectively divides the facial region into seven anatomically significant areas, i.e., left eye, right eye, nose, lips, left cheek, right cheek, and chin, enabling precise localization of muscle movements associated with micro-expressions. Optical flow-based motion analysis is employed to extract temporal changes in facial pixels across video sequences, ensuring the capture of subtle dynamic patterns. To enhance spatial feature extraction, a graph-based connectivity matrix models relationships between facial regions, which are then processed using GCNs to learn spatial dependencies and muscle activation patterns. Feature centering and L2 normalization are applied to improve feature robustness and ensure consistent representation across varying datasets. Additionally, a novel metric learning-based prototype classifier is introduced to mitigate class imbalance by incorporating a loss function that penalizes misclassifications of minority classes, thereby improving recognition performance on underrepresented categories. The proposed method achieves state-of-the-art performance with recognition accuracies of 81.83% on the SMIC dataset, 87.83% on CASME, and 93.48% on CASME II, significantly surpassing existing methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced Micro-Expression Recognition Through Graph Convolutional Networks and Metric Learning

  • Sreenivasu Bhukya,
  • L. Nirmala Devi,
  • A. Nageswar Rao

摘要

Micro-expression recognition (MER) remains a challenging task due to the subtle and transient nature of micro-expressions, which often results in poor generalization and low recognition accuracy in existing approaches. This research introduces a novel MER framework that integrates graph convolutional networks (GCNs) with facial action unit features to address critical limitations, including ineffective feature extraction, class imbalance, and insufficient temporal modeling. The proposed method effectively divides the facial region into seven anatomically significant areas, i.e., left eye, right eye, nose, lips, left cheek, right cheek, and chin, enabling precise localization of muscle movements associated with micro-expressions. Optical flow-based motion analysis is employed to extract temporal changes in facial pixels across video sequences, ensuring the capture of subtle dynamic patterns. To enhance spatial feature extraction, a graph-based connectivity matrix models relationships between facial regions, which are then processed using GCNs to learn spatial dependencies and muscle activation patterns. Feature centering and L2 normalization are applied to improve feature robustness and ensure consistent representation across varying datasets. Additionally, a novel metric learning-based prototype classifier is introduced to mitigate class imbalance by incorporating a loss function that penalizes misclassifications of minority classes, thereby improving recognition performance on underrepresented categories. The proposed method achieves state-of-the-art performance with recognition accuracies of 81.83% on the SMIC dataset, 87.83% on CASME, and 93.48% on CASME II, significantly surpassing existing methods.