<p>Multimodal Emotion Recognition in Conversation (MERC) aims to predict the emotional state of each utterance in a dialogue by leveraging contextual cues and multiple modalities. Consequently, the effective modeling of multimodal contextual information is central to this task. However, existing MERC approaches often conflate the modeling of intra- and inter-modal interactions across utterances and fail to effectively model the global comprehensive interactions among all modalities, thereby limiting the accuracy of MERC. To address these problems, we propose a novel dual-branch multimodal fusion network based on graph and attention (DFGAnet) for MERC, which primarily comprises two parallel branches, namely the unimodal branch multi graph propagation network (UMGPN) and the cross-modal branch multi attention interaction network (CMAIN). UMGPN focuses on adaptively modeling intra-modal interactions across utterances to extract diverse intra-modal contextual information. Meanwhile, CMAIN is specialized in modeling the global comprehensive interactions among all modalities to sufficiently extract inter-modal complementary information. Such a parallel learning architecture alleviates fusion conflicts and information redundancy, while preserving both unimodal and cross-modal features at the final fusion stage, thus promoting comprehensive multimodal integration. Furthermore, an auxiliary cross-modal fusion loss (ACFL) is introduced to further promote cross-modal representation learning. Finally, the IEMOCAP and MELD datasets are adopted to estimate the model performance and results indicate that the proposed method outperforms the comparative methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DFGAnet: a dual-branch multimodal fusion network based on graph and attention for emotion recognition in conversation

  • Wenzhuo Liu,
  • Taoying Li,
  • Yijia Chen

摘要

Multimodal Emotion Recognition in Conversation (MERC) aims to predict the emotional state of each utterance in a dialogue by leveraging contextual cues and multiple modalities. Consequently, the effective modeling of multimodal contextual information is central to this task. However, existing MERC approaches often conflate the modeling of intra- and inter-modal interactions across utterances and fail to effectively model the global comprehensive interactions among all modalities, thereby limiting the accuracy of MERC. To address these problems, we propose a novel dual-branch multimodal fusion network based on graph and attention (DFGAnet) for MERC, which primarily comprises two parallel branches, namely the unimodal branch multi graph propagation network (UMGPN) and the cross-modal branch multi attention interaction network (CMAIN). UMGPN focuses on adaptively modeling intra-modal interactions across utterances to extract diverse intra-modal contextual information. Meanwhile, CMAIN is specialized in modeling the global comprehensive interactions among all modalities to sufficiently extract inter-modal complementary information. Such a parallel learning architecture alleviates fusion conflicts and information redundancy, while preserving both unimodal and cross-modal features at the final fusion stage, thus promoting comprehensive multimodal integration. Furthermore, an auxiliary cross-modal fusion loss (ACFL) is introduced to further promote cross-modal representation learning. Finally, the IEMOCAP and MELD datasets are adopted to estimate the model performance and results indicate that the proposed method outperforms the comparative methods.