<p>Artificial intelligence-based techniques have become influential in depression recognition, offering objective solutions to a complex challenge. However, many deep learning methodologies lack effective mechanisms for emphasizing key features from different regions of the face, leading to suboptimal performance by failing to capture high-level semantic features. Therefore, to address these gaps, a multi-channel neural network (MCNN) with integrated channel-wise attention (CA) mechanism across multiple branches and feature fusion is proposed in this work. The integrated CA mechanism plays a vital role in effectively capturing important features from different regions of the face, thereby improving the model’s ability to recognize depression. With this incorporation, the residual network with 50 layers is transformed into a feature extraction module to effectively exploit global (facial) and local (mouth and eyes) features. Experiments are conducted on the Audio-Visual Emotion Challenge 2014 and Changzhou No. 2 People’s Hospital datasets. The MCNN achieved commendable results, notably a mean absolute error (MAE) of 8.04 and root mean square error (RMSE) of 9.65 on the former, whereas an MAE of 6.84 and RMSE of 8.77 on the latter. Contributions from this work underscore a novel direction and establish a benchmark for advancing depression recognition methodologies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MCNN: Multi-channel neural network with channel-wise attention for facial Expression-Based depression recognition

  • Muhammad Turyalai Khan,
  • Muhammad Imran,
  • Maria Kanwal

摘要

Artificial intelligence-based techniques have become influential in depression recognition, offering objective solutions to a complex challenge. However, many deep learning methodologies lack effective mechanisms for emphasizing key features from different regions of the face, leading to suboptimal performance by failing to capture high-level semantic features. Therefore, to address these gaps, a multi-channel neural network (MCNN) with integrated channel-wise attention (CA) mechanism across multiple branches and feature fusion is proposed in this work. The integrated CA mechanism plays a vital role in effectively capturing important features from different regions of the face, thereby improving the model’s ability to recognize depression. With this incorporation, the residual network with 50 layers is transformed into a feature extraction module to effectively exploit global (facial) and local (mouth and eyes) features. Experiments are conducted on the Audio-Visual Emotion Challenge 2014 and Changzhou No. 2 People’s Hospital datasets. The MCNN achieved commendable results, notably a mean absolute error (MAE) of 8.04 and root mean square error (RMSE) of 9.65 on the former, whereas an MAE of 6.84 and RMSE of 8.77 on the latter. Contributions from this work underscore a novel direction and establish a benchmark for advancing depression recognition methodologies.