The study presents a comparative evaluation of five Convolutional Neural Network (CNN) architectures – ResNet34v2, Xception, EfficientNetB2, MobileNetV3Large, and PAtt-Lite – for facial emotion recognition (FER) using three diverse datasets: FERPlus, CK+, and AffectNet. The primary contribution lies in cross-dataset evaluation using FERPlus, CK+, and AffectNet, addressing the critical challenge of model generalization to unseen data. Experimental results reveal that PAtt-Lite achieves the highest accuracy (96.2%) and balanced accuracy (68.4%) on FERPlus due to its specialized design. ResNet34v2 records the best accuracy (83.6%) on CK+, though its balanced accuracy is lower, highlighting challenges with underrepresented emotions. On AffectNet, all models performed much worse. PAtt-Lite model has difficulty with recognizing emotions like contempt and disgust, reflecting the complexity and class imbalance. Despite these variations, statistical analysis using the Friedman test indicates no significant differences in balanced accuracy and F-measure across models. These findings underscore that dataset characteristics, such as size, diversity, and label quality, play a crucial role in model performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Analysis of Convolutional Neural Network Architectures for Emotion Recognition in Facial Images

  • Małgorzata Przybyła-Kasperek,
  • Krystian Zieliński

摘要

The study presents a comparative evaluation of five Convolutional Neural Network (CNN) architectures – ResNet34v2, Xception, EfficientNetB2, MobileNetV3Large, and PAtt-Lite – for facial emotion recognition (FER) using three diverse datasets: FERPlus, CK+, and AffectNet. The primary contribution lies in cross-dataset evaluation using FERPlus, CK+, and AffectNet, addressing the critical challenge of model generalization to unseen data. Experimental results reveal that PAtt-Lite achieves the highest accuracy (96.2%) and balanced accuracy (68.4%) on FERPlus due to its specialized design. ResNet34v2 records the best accuracy (83.6%) on CK+, though its balanced accuracy is lower, highlighting challenges with underrepresented emotions. On AffectNet, all models performed much worse. PAtt-Lite model has difficulty with recognizing emotions like contempt and disgust, reflecting the complexity and class imbalance. Despite these variations, statistical analysis using the Friedman test indicates no significant differences in balanced accuracy and F-measure across models. These findings underscore that dataset characteristics, such as size, diversity, and label quality, play a crucial role in model performance.