This research explores the development of a deep learning-based audio-visual emotion recognition system, aiming to enhance the accuracy and robustness of emotion classification by integrating multiple modalities. Traditional speech emotion recognition (SER) systems often rely on unimodal data, which limits their ability to fully capture human emotional expressions. Our study leverages the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) to implement a multimodal approach, combining audio and visual data. The proposed model incorporates Convolutional Neural Networks (CNNs), Long Short-Term Memory (LSTM) networks, and attention mechanisms to improve performance. Experimental results demonstrate that the attention-based audio model achieves the highest accuracy of 62%, outperforming other tested configurations. The study highlights the potential of integrating attention mechanisms and multimodal data to enhance SER systems, while also identifying areas for future research, such as utilizing additional datasets and transfer learning techniques to further improve model performance and generalizability.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Audio-Visual Emotion Recognition Using Deep Learning Methods

  • Mukhambet Tolegenov,
  • Lakshmi Babu Saheer,
  • Mahdi Maktabdar Oghaz

摘要

This research explores the development of a deep learning-based audio-visual emotion recognition system, aiming to enhance the accuracy and robustness of emotion classification by integrating multiple modalities. Traditional speech emotion recognition (SER) systems often rely on unimodal data, which limits their ability to fully capture human emotional expressions. Our study leverages the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) to implement a multimodal approach, combining audio and visual data. The proposed model incorporates Convolutional Neural Networks (CNNs), Long Short-Term Memory (LSTM) networks, and attention mechanisms to improve performance. Experimental results demonstrate that the attention-based audio model achieves the highest accuracy of 62%, outperforming other tested configurations. The study highlights the potential of integrating attention mechanisms and multimodal data to enhance SER systems, while also identifying areas for future research, such as utilizing additional datasets and transfer learning techniques to further improve model performance and generalizability.