Imbalance-Aware Multimodal Attention-Enabled Latent Space Oversampling for Emotion Recognition
摘要
Recent achievements in deep learning have provided a thrust in the development of emotion recognition systems. Researchers have developed multimodal emotion recognition models that leverage both facial and non-facial data such as text and audio. However, most of the models suffer from the well-known problem of imbalanced classes which occurs due to the unavailability of data in many emotion categories. Motivated by this, the current article has proposed an imbalance-aware multimodal attention-enabled latent space oversampling framework to address the uneven distribution of data samples in various emotion categories. The primary contributions of the manuscript are as follows: (i) A multimodal deep learning framework is developed to predict emotion from facial images and textual data. (ii) The model includes a separate attention-enabled CNN-based image encoder to extract a latent representation of input image. (iii) A separate BERT encoder is used to extract textual features. (iv) Both latent vectors are fused to obtain a multimodal latent representation of both image and text data. Thereafter, the latent space oversampling method is used to address the class imbalance problem. The balanced latent vectors are then employed to train shallow learning models to predict the emotion. Experiments have revealed that the attention-enabled latent space oversampling framework can effectively predict emotion from multimodal data with greater accuracy than multimodal models without attention and latent space oversampling.