<p>Emotion recognition from facial images is an important area of computer vision with applications in healthcare, human–computer interaction and entertainment. This paper proposes an Advanced Emotion Recognition from Facial Images and Speech Signals using Complex-Value Spatio-Temporal Graph Convolutional Neural Network (AER-FISS-CSGCNN). The Ryerson audio-visual dataset of emotional speech and song is used for training. Visual data is normalized using Kernel-based Ensemble Gaussian Mixture Filtering, while audio noise is reduced using Strong Tracking Variational Bayesian Adaptive Kalman Filter. Feature extraction is performed using dual vision transformer for texture features like Entropy, Contrast and Correlation and Bayesian Weighted Random Forest for spectral audio features like MFCCs. a multi-tensor fusion network integrates both modalities. The fused features are classified under Complex-Value Spatio-Temporal Graph Convolutional Neural Network (CSGCNN) with weight optimization utilizing red-billed blue magpie optimizer. The system classifies emotions, such as happy, sad, angry, surprised. The experimental results show that the AER-FISS-CSGCNN outperforms existing methods by achieving up to 31.66% higher accuracy and 29.50% higher precision.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advanced emotion recognition from facial images and speech signals using complex-value spatio-temporal graph convolutional neural networks

  • Shine P. Xavier,
  • Saju P. John

摘要

Emotion recognition from facial images is an important area of computer vision with applications in healthcare, human–computer interaction and entertainment. This paper proposes an Advanced Emotion Recognition from Facial Images and Speech Signals using Complex-Value Spatio-Temporal Graph Convolutional Neural Network (AER-FISS-CSGCNN). The Ryerson audio-visual dataset of emotional speech and song is used for training. Visual data is normalized using Kernel-based Ensemble Gaussian Mixture Filtering, while audio noise is reduced using Strong Tracking Variational Bayesian Adaptive Kalman Filter. Feature extraction is performed using dual vision transformer for texture features like Entropy, Contrast and Correlation and Bayesian Weighted Random Forest for spectral audio features like MFCCs. a multi-tensor fusion network integrates both modalities. The fused features are classified under Complex-Value Spatio-Temporal Graph Convolutional Neural Network (CSGCNN) with weight optimization utilizing red-billed blue magpie optimizer. The system classifies emotions, such as happy, sad, angry, surprised. The experimental results show that the AER-FISS-CSGCNN outperforms existing methods by achieving up to 31.66% higher accuracy and 29.50% higher precision.