This paper introduces a robust and efficient multimodal deep learning network for human emotion recognition that integrates audio and visual data, among Indian population. The proposed approach is applied on custom built audio-visual dataset consisting of 122 videos of 61 Indian participants in the age group of 18–21 years, of which 36 were male and 25 were female, capturing a wide spectrum of emotional expressions across Indian population. The core of our study revolves around a hybrid machine learning-deep learning network, systematically evaluated at multiple stages of development. A key achievement is the introduction of a multimodal model that synthesizes audio and visual modalities, providing a holistic understanding of emotional expressions. Machine Learning algorithms such as XGBoost, and DCNN models such as ResNet50, were tested for audio and visual modality simulation respectively. Our results demonstrate the effectiveness of this approach, with the best multimodal model achieving a training accuracy of 0.9143 and a validation accuracy of 0.8489. We also explore architectural variations, revealing potential improvements. This work has the potential to significantly enhance human–computer interaction, affective computing, and sentiment analysis, providing a more comprehensive understanding of culturally diverse human emotions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Multimodal Deep Learning Approach for Emotion Recognition in a Diverse Indian Cultural Context

  • Ruhina Karani,
  • Vijay Harkare,
  • Krishna Kamath,
  • Khushi Gupta,
  • Om Shukla,
  • Sharmishta Desai

摘要

This paper introduces a robust and efficient multimodal deep learning network for human emotion recognition that integrates audio and visual data, among Indian population. The proposed approach is applied on custom built audio-visual dataset consisting of 122 videos of 61 Indian participants in the age group of 18–21 years, of which 36 were male and 25 were female, capturing a wide spectrum of emotional expressions across Indian population. The core of our study revolves around a hybrid machine learning-deep learning network, systematically evaluated at multiple stages of development. A key achievement is the introduction of a multimodal model that synthesizes audio and visual modalities, providing a holistic understanding of emotional expressions. Machine Learning algorithms such as XGBoost, and DCNN models such as ResNet50, were tested for audio and visual modality simulation respectively. Our results demonstrate the effectiveness of this approach, with the best multimodal model achieving a training accuracy of 0.9143 and a validation accuracy of 0.8489. We also explore architectural variations, revealing potential improvements. This work has the potential to significantly enhance human–computer interaction, affective computing, and sentiment analysis, providing a more comprehensive understanding of culturally diverse human emotions.