Fusing facial and speech cues for enhanced multimodal emotion recognition
摘要
Emotion recognition is a technology that enables computers to recognize and interpret human emotions by analyzing facial expressions, voice, text or physiological signals. It finds applications in human–computer interaction, mental health assessment, and personalized content recommendation, offering insights into user sentiment and engagement. In this paper, we introduced an innovative approach to emotion recognition which combines facial expressions and speech cues within a multimodal system. This fusion of two distinct modalities is achieved through two specific methods: feature-level fusion and decision-level fusion. To evaluate the effectiveness of our approach, we conduct experiments using the eNTERFACE'05 dataset. Our comparative analysis reveals that the integration of fusion-based techniques can substantially enhance the performance of emotion recognition systems. Additionally, our findings highlight the superiority of feature-level fusion over decision-level fusion in terms of overall performance.