Audio-visual expression-based emotion recognition model for neglected people in real-time: a late-fusion approach
摘要
Analyzing emotions extracted from diverse media sources like video and audio presents a significant challenge in identifying human mental health indicators. This is particularly crucial in studying the mental well-being of Transgender communities, who often encounter difficulty in accessing emergency services due to discrimination. This study employed Speech Emotion Recognition (SER) for identifying emotions in audio signals and a dynamic Facial Expression Recognition (FER) mechanism for analyzing emotions in video frames. The SER model utilizes deep learning techniques and Explainable AI (XAI) to optimize the audio feature vector, achieving a remarkable accuracy on the RAVDESS dataset. Benchmark metrics employed to evaluate the performance of the SER model, shows that our employed SER surpasses previously engaged Machine Learning (ML) models. Similarly, the FER model (VGG-Face), based on transfer learning proved to be efficient in analyzing YouTube video frames maintaining high accuracy. Both SER and FER models are then applied to assess the mental health of transgender individuals. To consolidate our findings, we adopt a late fusion-based approach to combine classified emotions, employing various voting methods such as primary, Condorcet, non-Condorcet Preferential, and valence-arousal. The SER method reveals that anger, sadness, and disgust are dominant emotions, while the FER method indicates an equal presence of negativity and neutrality. Our consolidate results based on late fusion approaches highlighted the dominance of anger emotion. In essence, this study offers valuable insights into the mental well-being of marginalized communities, providing a foundation for the development of tailored programs to support their mental health needs.