A cross-modal fusion network based on dual attention mechanism for emotion recognition in conversation
摘要
Emotion recognition in conversations (ERC) plays a vital role in domains such as empathic dialog systems. However, existing methods often fail to utilize the complementary nature of multimodal information and assign equal importance to each modality, limiting their ability to capture key emotional cues. To address these challenges, this paper proposes CFDAN-ERC, a Cross-modal Fusion Network based on a Dual Attention Mechanism for Emotion Recognition in Conversations. CFDAN-ERC employs differentiated fusion strategies for textual and video-audio modalities, emphasizing the central role of textual information while leveraging visual and acoustic modalities as auxiliary sources. The model comprises two core modules: sLSTM-based Uni-Modality Encoding (sLSTM-UME) to extract contextual emotional cues and mitigate inter-modal heterogeneity. And the DA-CME module achieves comprehensive multimodal modeling by extracting global–local affective cues within modalities and capturing complementary information between modalities through the Efficient Multiscale External Attention (EMEA) mechanism and the multi-head attention mechanism proposed in this paper, respectively. Furthermore, Soft-HGR Loss and Poly Focal Loss are introduced to enhance multimodal fusion and ease challenges in recognizing minority and semantically similar emotion categories. Extensive experimental results demonstrate that CFDAN-ERC outperforms existing state-of-the-art methods in sentiment classification on the two benchmark ERC datasets, MELD and IEMOCAP. Notably, it achieves significant improvements in minority sentiment categories and semantically similar emotion categories.