GCIF: graph based cross-modal information fusion for conversational emotion recognition
摘要
Multimodal Emotion Recognition in Conversations (MERC) represents a pivotal research avenue within the realms of human-computer interaction and affective computing. The fundamental challenge is to develop effective models for multimodal contextual information and to integrate complementary multimodal data. Given the superior performance of Graph Neural Networks (GNNs) in relation modelling, this paper puts forward a Graph Based Cross-modal Information Fusion (GCIF) for conversational emotion recognition to tackle the problems of redundant information generation, information loss and over-smoothing that have been identified in the existing GNN methods for multimodal information fusion. The GCIF builds a graph network with a configurable fixed context window and integrates speaker information for multimodal data. The generality space and individuality space are created in order to maintain the consistency and specificity of multimodal features. GCIF employs a module based on the improved graph attention networks to achieve pairwise fusion of multimodal information, thereby reducing the difficulty of multimodal fusion and effectively alleviating the problems of heterogeneity and over-smoothing. The experimental results on the two public benchmark datasets demonstrate that our GCIF can effectively promote performance for MERC. Furthermore, this paper conducts extensive experiments to discuss and analyze the impact of different settings on the performance of GCIF, thereby validating the effectiveness of the model.