Emotion Recognition in Conversations (ERC) is a key task in Human-Computer Interaction (HCI), yet challenges in modeling emotion shifts and context dependencies hinder further improvements in classification accuracy. To address the challenges, we innovatively propose an emotion recognition architecture, View-Fusion Dual-Stream Framework (VF-Dual), skillfully integrating a multi-view decoder, a dual-stream context extraction module, and an adaptive decoupled contrastive learning strategy. Specifically, we design a multi-view decoder to capture optimistic and pessimistic emotions, ensuring that both views’ emotions are preserved and complement each other. Subsequently, these encoded emotion representations are combined with speaker information and fed into the dual-stream information extraction module, significantly enhancing the model’s ability to capture context dependencies by deeply mining the complex correlations among different ranges of context. Furthermore, we raise the adaptive decoupled contrastive learning to avoid mutual interference between positive and negative samples, ensuring the purity and differentiation of emotion representations. Additionally, a momentum contrast mechanism is introduced to increase the number and diversity of negative samples. Extensive experimental results demonstrate that our proposed VF-Dual significantly outperforms existing state-of-the-art models in ERC task.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VF-Dual: View-Fusion Dual-Stream Architecture for Emotion Recognition in Conversations

  • Feifei Xu,
  • Qinghan Du,
  • Xiaoyan Yu

摘要

Emotion Recognition in Conversations (ERC) is a key task in Human-Computer Interaction (HCI), yet challenges in modeling emotion shifts and context dependencies hinder further improvements in classification accuracy. To address the challenges, we innovatively propose an emotion recognition architecture, View-Fusion Dual-Stream Framework (VF-Dual), skillfully integrating a multi-view decoder, a dual-stream context extraction module, and an adaptive decoupled contrastive learning strategy. Specifically, we design a multi-view decoder to capture optimistic and pessimistic emotions, ensuring that both views’ emotions are preserved and complement each other. Subsequently, these encoded emotion representations are combined with speaker information and fed into the dual-stream information extraction module, significantly enhancing the model’s ability to capture context dependencies by deeply mining the complex correlations among different ranges of context. Furthermore, we raise the adaptive decoupled contrastive learning to avoid mutual interference between positive and negative samples, ensuring the purity and differentiation of emotion representations. Additionally, a momentum contrast mechanism is introduced to increase the number and diversity of negative samples. Extensive experimental results demonstrate that our proposed VF-Dual significantly outperforms existing state-of-the-art models in ERC task.