<p>Emotion recognition in conversations (ERC) plays a vital role in domains such as empathic dialog systems. However, existing methods often fail to utilize the complementary nature of multimodal information and assign equal importance to each modality, limiting their ability to capture key emotional cues. To address these challenges, this paper proposes CFDAN-ERC, a Cross-modal Fusion Network based on a Dual Attention Mechanism for Emotion Recognition in Conversations. CFDAN-ERC employs differentiated fusion strategies for textual and video-audio modalities, emphasizing the central role of textual information while leveraging visual and acoustic modalities as auxiliary sources. The model comprises two core modules: sLSTM-based Uni-Modality Encoding (sLSTM-UME) to extract contextual emotional cues and mitigate inter-modal heterogeneity. And the DA-CME module achieves comprehensive multimodal modeling by extracting global–local affective cues within modalities and capturing complementary information between modalities through the Efficient Multiscale External Attention (EMEA) mechanism and the multi-head attention mechanism proposed in this paper, respectively. Furthermore, Soft-HGR Loss and Poly Focal Loss are introduced to enhance multimodal fusion and ease challenges in recognizing minority and semantically similar emotion categories. Extensive experimental results demonstrate that CFDAN-ERC outperforms existing state-of-the-art methods in sentiment classification on the two benchmark ERC datasets, MELD and IEMOCAP. Notably, it achieves significant improvements in minority sentiment categories and semantically similar emotion categories.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A cross-modal fusion network based on dual attention mechanism for emotion recognition in conversation

  • Xinheng Wang,
  • Lun Xie,
  • Chiqin Li,
  • Mengsheng Wang,
  • Ziyang Liu,
  • Xiaolan Peng,
  • Zhiliang Wang

摘要

Emotion recognition in conversations (ERC) plays a vital role in domains such as empathic dialog systems. However, existing methods often fail to utilize the complementary nature of multimodal information and assign equal importance to each modality, limiting their ability to capture key emotional cues. To address these challenges, this paper proposes CFDAN-ERC, a Cross-modal Fusion Network based on a Dual Attention Mechanism for Emotion Recognition in Conversations. CFDAN-ERC employs differentiated fusion strategies for textual and video-audio modalities, emphasizing the central role of textual information while leveraging visual and acoustic modalities as auxiliary sources. The model comprises two core modules: sLSTM-based Uni-Modality Encoding (sLSTM-UME) to extract contextual emotional cues and mitigate inter-modal heterogeneity. And the DA-CME module achieves comprehensive multimodal modeling by extracting global–local affective cues within modalities and capturing complementary information between modalities through the Efficient Multiscale External Attention (EMEA) mechanism and the multi-head attention mechanism proposed in this paper, respectively. Furthermore, Soft-HGR Loss and Poly Focal Loss are introduced to enhance multimodal fusion and ease challenges in recognizing minority and semantically similar emotion categories. Extensive experimental results demonstrate that CFDAN-ERC outperforms existing state-of-the-art methods in sentiment classification on the two benchmark ERC datasets, MELD and IEMOCAP. Notably, it achieves significant improvements in minority sentiment categories and semantically similar emotion categories.