错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Synergizing Senses: Advancing Multimodal Emotion Recognition in Human-Computer Interaction with MFF-CNN

  • Kamal Upreti,
  • Prashant Vats,
  • Khushboo Malik,
  • Rajesh Verma,
  • Prakash Divakaran,
  • Divya Gangwar

摘要

Optimizing the authenticity and efficacy of interactions between humans and computers is largely dependent on emotion detection. The MFF-CNN framework is used in this work to present a unique method for multidimensional emotion identification. The MFF-CNN model is a combination of approaches that combines convolutional neural networks and multimodal fusion. It is intended to efficiently collect and integrate data from several modalities, including spoken words and human facial expressions. The first step in the suggested system's implementation is gathering a multimodal dataset with emotional labels added to it. The MFF-CNN receives input features in the form of retrieved facial landmarks and voice signal spectroscopy reconstructions. Convolutional layers are used by the model to understand hierarchies’ spatial and temporal structures, which improves its capacity to recognize complex emotional signals. Our experimental assessment shows that the MFF-CNN outperforms conventional unimodal emotion recognition algorithms. Improved preciseness, reliability, and adaptability across a range of emotional states are the outcomes of fusing the linguistic and face senses. Additionally, visualization methods improve the interpretability of the model and offer insights into the learnt representations. By providing a practical and understandable method for multimodal emotion identification, this study advances the field of human-computer interaction. The MFF-CNN architecture opens the door to more organic and psychologically understanding human-computer interactions by showcasing its possibilities for practical applications.