<p>Emotion recognition plays a vital role in human-computer interaction (HCI), enabling systems to interpret and respond appropriately to human emotions. While traditional methods often rely on either speech or facial expressions alone, human emotions are naturally conveyed through both modalities simultaneously. This study proposes a CNN-BiLSTM-based multimodal framework that captures spatial features from Mel-spectrograms and facial expressions, and temporal or sequential features from MFCC for robust emotion detection. For vocal emotion analysis, BiLSTM is applied to MFCCs to extract sequential features, while CNN is used for spatial feature extraction from both Mel-spectrograms and facial images. Then feature-level fusion is used to concatenate the extracted features to recognize the emotion from both vocal and facial expressions. The “feature-level fusion” approach means that CNN and BiLSTM process their respective modalities separately to extract features, and their outputs are fused at a later stage before making the final decision. The model is trained and tested on a merged corpus from RAVDESS, CREMA-D, and SAVEE, comprising actors from diverse ethnic backgrounds and emotional styles. This provides a more realistic and generalizable evaluation setting compared to prior works trained on single, often homogeneous datasets. The proposed model achieves an accuracy of 86.93%, outperforming unimodal approaches and demonstrating strong performance in recognizing emotions such as happiness and anger. The novelty lies in the three-stream (facial images, MFCCs, and Mel-spectrograms) feature-level fusion architecture that integrates two spatial and one temporal feature sets, enhancing emotion recognition capabilities in real-world applications. This work has implications for affective computing, intelligent virtual assistants, surveillance systems, and customer experience analysis.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Three-Stream Feature-Level Fusion Based Multimodal Emotion Recognition from Vocal and Facial Expressions for Human-Computer Interaction

  • Samiddha Chakrabarti,
  • Parthasarathi De

摘要

Emotion recognition plays a vital role in human-computer interaction (HCI), enabling systems to interpret and respond appropriately to human emotions. While traditional methods often rely on either speech or facial expressions alone, human emotions are naturally conveyed through both modalities simultaneously. This study proposes a CNN-BiLSTM-based multimodal framework that captures spatial features from Mel-spectrograms and facial expressions, and temporal or sequential features from MFCC for robust emotion detection. For vocal emotion analysis, BiLSTM is applied to MFCCs to extract sequential features, while CNN is used for spatial feature extraction from both Mel-spectrograms and facial images. Then feature-level fusion is used to concatenate the extracted features to recognize the emotion from both vocal and facial expressions. The “feature-level fusion” approach means that CNN and BiLSTM process their respective modalities separately to extract features, and their outputs are fused at a later stage before making the final decision. The model is trained and tested on a merged corpus from RAVDESS, CREMA-D, and SAVEE, comprising actors from diverse ethnic backgrounds and emotional styles. This provides a more realistic and generalizable evaluation setting compared to prior works trained on single, often homogeneous datasets. The proposed model achieves an accuracy of 86.93%, outperforming unimodal approaches and demonstrating strong performance in recognizing emotions such as happiness and anger. The novelty lies in the three-stream (facial images, MFCCs, and Mel-spectrograms) feature-level fusion architecture that integrates two spatial and one temporal feature sets, enhancing emotion recognition capabilities in real-world applications. This work has implications for affective computing, intelligent virtual assistants, surveillance systems, and customer experience analysis.