Facial Expression Recognition (FER), as an important field of research within computer vision, holds substantial application value. However, numerous studies have revealed critical challenges in FER, including high inter-class similarity, large intra-class discrepancy, and scale sensitivity, which collectively contribute to expression uncertainty and may even induce annotation errors. Although some recent approaches have made improvements, they have not tackled these challenges concurrently. This study introduces a new FER model called DSF-SL (Dual-Stream Fusion with Soft Labels). First, we develop an MSCF (Multi-Scale Cross-Fusion) module to guide expression features using landmark information and dynamically extract multi-scale features. Then, we introduce a HA (Hierarchical Attention) mechanism to compute both channel and sample attention, enabling the model to capture unique features for each expression class. Finally, we implement an SLS (Soft Label Smoothing) strategy to dynamically refine hard labels into soft labels. Extensive experiments demonstrate that DSF-SL achieves state-of-the-art performance on multiple benchmark datasets: RAF-DB (92.96%), AffectNet-7cls (72.60%), and AffectNet-8cls (65.83%), along with a competitive accuracy on FERPlus (89.95%). These results substantiate that DSF-SL exhibits remarkable superiority, demonstrating significant potential to advance the development of FER technologies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DSF-SL: A Dual-Stream Fusion Network with Soft Labels for Robust Facial Expression Recognition

  • Haoyu Liu,
  • Xiaoqing Jiang,
  • Chenyang Liang,
  • Peizhi Sun,
  • Jianwei Gu,
  • Bang Li,
  • Haoran Sun,
  • Tuchuan Chen,
  • Zhenxiang Chen

摘要

Facial Expression Recognition (FER), as an important field of research within computer vision, holds substantial application value. However, numerous studies have revealed critical challenges in FER, including high inter-class similarity, large intra-class discrepancy, and scale sensitivity, which collectively contribute to expression uncertainty and may even induce annotation errors. Although some recent approaches have made improvements, they have not tackled these challenges concurrently. This study introduces a new FER model called DSF-SL (Dual-Stream Fusion with Soft Labels). First, we develop an MSCF (Multi-Scale Cross-Fusion) module to guide expression features using landmark information and dynamically extract multi-scale features. Then, we introduce a HA (Hierarchical Attention) mechanism to compute both channel and sample attention, enabling the model to capture unique features for each expression class. Finally, we implement an SLS (Soft Label Smoothing) strategy to dynamically refine hard labels into soft labels. Extensive experiments demonstrate that DSF-SL achieves state-of-the-art performance on multiple benchmark datasets: RAF-DB (92.96%), AffectNet-7cls (72.60%), and AffectNet-8cls (65.83%), along with a competitive accuracy on FERPlus (89.95%). These results substantiate that DSF-SL exhibits remarkable superiority, demonstrating significant potential to advance the development of FER technologies.