<p>Facial expressions are an essential part of nonverbal communication, and have the ability to express compound expressions, which are combinations of dominant and complementary emotions. Recent research for classifying compound expressions are limited to occlusion, inter-class similarities, and misclassification, thus leading to performance degradation in real-world environments. To address these challenges, this paper proposes a novel TriphaseNet, consisting of a three-stage framework. The initial stage comprises a novel Depth Attentive Couple-Trans module (DACT) to obtain fine-grained information and capture long-range dependencies, Multi-Kernel Pattern Extractor Module (MPEM) provides global and contextual cues and Channel-Aware Excitation Module (CAEM) to focus on the most enlightening aspects, which are complex for differentiating compound emotions. In the second stage, a novel interleaved feature stacking technique maintains spatial correlations and finds the predominant areas from different feature maps. In the third stage, a 2D facial landmark mask is generated to highlight the facial structure in the facial image and ensures a more robust representation of expressions against pose variations and occlusions. Rigorous experiments on RAF-DB and CFEE datasets show that the proposed model significantly outperforms well, achieving 85.76% and 71.53% accuracy on basic and compound RAF-DB, 90.53% and 79.16% on basic and compound CFEE datasets. These results emphasize the model’s capability to handle the above challenges in the facial images captured in natural scenarios for compound expressions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Triphasenet: Interleaved feature stacking with a deep extractor using couple-trans for compound facial expression recognition

  • Vidhya V,
  • G. Devasena

摘要

Facial expressions are an essential part of nonverbal communication, and have the ability to express compound expressions, which are combinations of dominant and complementary emotions. Recent research for classifying compound expressions are limited to occlusion, inter-class similarities, and misclassification, thus leading to performance degradation in real-world environments. To address these challenges, this paper proposes a novel TriphaseNet, consisting of a three-stage framework. The initial stage comprises a novel Depth Attentive Couple-Trans module (DACT) to obtain fine-grained information and capture long-range dependencies, Multi-Kernel Pattern Extractor Module (MPEM) provides global and contextual cues and Channel-Aware Excitation Module (CAEM) to focus on the most enlightening aspects, which are complex for differentiating compound emotions. In the second stage, a novel interleaved feature stacking technique maintains spatial correlations and finds the predominant areas from different feature maps. In the third stage, a 2D facial landmark mask is generated to highlight the facial structure in the facial image and ensures a more robust representation of expressions against pose variations and occlusions. Rigorous experiments on RAF-DB and CFEE datasets show that the proposed model significantly outperforms well, achieving 85.76% and 71.53% accuracy on basic and compound RAF-DB, 90.53% and 79.16% on basic and compound CFEE datasets. These results emphasize the model’s capability to handle the above challenges in the facial images captured in natural scenarios for compound expressions.