UCEMA: Uni-modal and cross-modal encoding network based on multi-head attention for emotion recognition in conversation
摘要
Emotion recognition in conversation (ERC) represents a pivotal research domain within affective computing, concentrating on discerning the emotional nuances embedded within individual utterances during conversational exchanges. The majority of current research focuses on modeling situational cues, with relatively little attention paid on affective tendencies inherent in emotional expression. Moreover, ensuring the fair representation of diverse modalities in emotional expression presents a significant challenge in effectively extracting synergies and insights from multi-modal data sources. To tackle these challenges, this study proposes a novel approach termed the Uni-Modal and Cross-Modal Encoding Network based on Multi-Head Attention (UCEMA) for ERC. The framework leverages two distinct encoding techniques, namely Uni-Modal Encoding based on Multi-Head Attention (UEMA) and Cross-Modal Encoding based on Multi-Head Attention (CEMA), to extract distinct emotional features from individual modes and facilitate the fusion of emotional attributes within the multi-modal context. Particular emphasis is placed on textual input as the primary mode of interaction. Additionally, this study employs Context Modeling (CM) to analyze the outcomes of emotion recognition in conversational contexts. A comprehensive comparative analysis of the UCEMA was conducted on two publicly available datasets, IEMOCAP and MELD. The results demonstrated that the UCEMA exhibited superior efficacy. It is noteworthy that the proposed framework effectively balances intra-modal emotional orientation information, inter-modal emotional association information, and context-related cues, thereby demonstrating superior performance in recognition accuracy compared to current state-of-the-art (SOTA) models.