错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Emotion Recognition Using Attention-Based Model with Language, Audio, and Video Modalities

  • Disha Sharma,
  • Manoj Jayabalan,
  • Nailya Sultanova,
  • Jamila Mustafina,
  • Danny Ngo Lung Yao

摘要

Multimodal emotion identification is becoming increasingly important in human–computer interaction due to the amount of emotional information in human communication. Multimodal emotion recognition is the technique of simultaneously considering several modalities to boost accuracy and robustness. As emotion identification studies become more vital to human–computer interactions, automatic emotion detection systems become increasingly necessary. However, a lack of data presents a problem for multimodal emotion identification. To address this issue, we suggest employing transfer learning, which uses pretrained models such as RoBERTa and attention-based mechanisms such as self-attention to extract relevant features from multiple modalities and multi-head attention to fuse data across modalities. The aim of this paper is to provide a strategy for reliably forecasting emotions in audio, visual, and text by merging and complementing aspects traditionally handled by humans with those typically handled by deep learning. During the study, three popular multimodal emotion recognition datasets, IEMOCAP, CMU-MOSI, and CMU-MOSEI, are analyzed and ranked based on their quality. This study will help in constructing the network with the right amount of focus placed on each feature modality by creating an architecture that efficiently combines textual characteristics retrieved from RoBERTa with other modality-based features. A model better than BERT is introduced as part of this work that helps to improve the performance.