A Multimodal Deep Learning Approach for Emotion Recognition in a Diverse Indian Cultural Context
摘要
This paper introduces a robust and efficient multimodal deep learning network for human emotion recognition that integrates audio and visual data, among Indian population. The proposed approach is applied on custom built audio-visual dataset consisting of 122 videos of 61 Indian participants in the age group of 18–21 years, of which 36 were male and 25 were female, capturing a wide spectrum of emotional expressions across Indian population. The core of our study revolves around a hybrid machine learning-deep learning network, systematically evaluated at multiple stages of development. A key achievement is the introduction of a multimodal model that synthesizes audio and visual modalities, providing a holistic understanding of emotional expressions. Machine Learning algorithms such as XGBoost, and DCNN models such as ResNet50, were tested for audio and visual modality simulation respectively. Our results demonstrate the effectiveness of this approach, with the best multimodal model achieving a training accuracy of 0.9143 and a validation accuracy of 0.8489. We also explore architectural variations, revealing potential improvements. This work has the potential to significantly enhance human–computer interaction, affective computing, and sentiment analysis, providing a more comprehensive understanding of culturally diverse human emotions.