Multimodal Attention CNN for Human Emotion Recognition
摘要
The human face is the mirror of the mind. The face generally tells all that is going on in one’s heart and mind. Just by looking at the faces of our known ones, we may easily guess their mood. But many times, when we meet some unfamiliar person, it’s hard to get his or her mood just by looking at their faces. This is just because the person may have a certain facial structure that makes them by default look angry, happy, or sad. So, we need to spend some time with that person to analyse other parameters before concluding their state of mood. The current work proposed a novel approach that integrates facial images with electroencephalography (EEG) signals for facial expression recognition tasks. When attention-based deep CNN analyses the facial traits of the subject, a parallel Long Short-Term Memory (LSTM) network analyses the EEG signals. A late fusion network combines the features extracted from both networks, and finally, a classification network tells about what is the current mood of the subject. Combining multiple modalities for emotion recognition has shown promising results when compared with other state-of-the-art models. There are multiple real-life applications of emotion recognition models, such as Advertisement Industry, Human–Robot Interaction, Automatic Depression Detection, Mood Audio/Video Players, etc.