Situation-Speaker Interactive Network with Joint Contrastive Learning for Few-Shot Emotion Recognition in Conversation
摘要
Few-shot emotion recognition in conversation (FSERC) aims to detect the emotions conveyed in each utterance of a conversation, relying solely on a limited number of training conversations. Existing works directly concatenate situation-aware information with speaker-aware information. However, these methods overlook the interaction between the two types of information. Additionally, query-centered methods suffer from inadequate prototype representations and class imbalance in conversations. In this paper, we propose a situation-speaker interactive network with joint contrastive learning (SSIJCL) for FSERC. Firstly, to attain rich interactive information, we design a situation-speaker interactive network (SSI). This module is to model the two types of information into a unified graph through triple-connection interaction. Secondly, to further obtain more separable and distinguishable prototype representations, we propose a joint contrastive learning (JCL) module. This module adopts a prototype-centered training strategy, which can be highly complementary to the query-centered one. Thirdly, to sample sparse classes in a balanced way, we propose an intra-conversation class-balanced sampler (ICS). This module ensures effective computation of each class prototype. Finally, experimental results on four benchmark datasets show that SSIJCL not only achieves state-of-the-art performance on FSERC, but also performs comparably to traditional supervised learning models.