Enhanced Recognition of Human Emotion via Multimodal Inputs Using Fusion LSTM + + Model
摘要
Emotion recognition research in human–computer interactions has necessitated the development of automatic emotion identification systems. People’s ability to identify their emotions may help them better manage their emotions and engage with different situations in their lives. Numerous studies have explored into emotion categorization methods, although majority of them simply consider one, a small number, or independent physiological signs. Recognizing the importance of distinct signals and their integration will enable the development of further informative, economical, and objective techniques for detecting emotions, processing, and interpretations. In this paper, a novel FusionLSTM + + is proposed that integrates multimodal inputs such as audio, text, and motion data, including facial expressions and hand movements by employing a hybrid neural network approach. It involves designing separate classification structures for each modality and fusing the outputs at the final layer, resulting in more reliable and accurate emotion detection. Experimental evaluation on the IEMOCAP dataset demonstrates the effectiveness of our neural network-based framework in capturing nuanced emotional expressions, contributing to the advancement of multimodal emotion recognition in human–computer interactions.