A Cross-Modal Correlation Fusion Network for Emotion Recognition in Conversations
摘要
The aim of Emotion Recognition in Conversations (ERC) is to predict the emotions conveyed in the utterances of a conversation. In this paper, we propose a Cross-Modal Correlation Fusion Network (CMCFN), which addresses the limitations of existing approaches to exploit correlations across multiple modalities and the difficulty of classifying tail emotion categories. The proposed Cross-Modal Correlation Encoder (CMCE) effectively models intricate cross-modal correlations in a conversation, facilitating efficient multimodal fusion. In addition, the designed Multimodal Contrastive Representation Learning Network (MCRLN) mitigates the difficulty in categorizing tail emotions by combining supervised contrastive learning and multimodal data augmentation. Experimental results on the IEMOCAP and MELD datasets demonstrate the effectiveness and superiority of our proposed CMCFN model.