Machine Learning Model Based on Deep Neural Networks for Emotion Detection Using Audio-Visual Modalities
摘要
Emotion identification is a challenging endeavor since emotions can manifest in so many different ways. Recent advances have been made in the field, particularly with the development of deep neural networks, which have shown to be astonishingly good at recognizing emotional states. This work set out to investigate the difficulties of model-level fusion in order to develop a sophisticated multimodal model capable of simultaneously assessing both video and audio inputs for Emotion Recognition. To rigorously evaluate the performance of this novel approach, the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) dataset was chosen as the ideal testing ground. This diverse dataset encompasses a wide range of emotional expressions, making it an excellent benchmark for assessing the proposed model’s accuracy and robustness. By combining cutting-edge deep learning techniques with a multimodal approach, this study aims to contribute to the evolving landscape of Emotion Recognition, with the potential to enhance applications in fields such as affective computing, human–computer interaction, and emotional well-being assessment.