HLPCMF: A Hypergraph Learning and Pairwise Cross-Modal Fusion Model for Multimodal Emotion Recognition in Conversation
摘要
Multimodal emotion recognition is critical for affective computing applications, but it faces growing privacy risks when handling sensitive user data, while generic models often fail to adapt to individual emotional characteristics. Existing methods rarely address both privacy protection and personalization simultaneously. This paper proposes HLPCMF, a novel multimodal emotion recognition model that integrates hypergraph learning with pairwise cross-modal fusion. Specifically, we construct multimodal hyperedges and temporal hyperedges to model high-order relationships among utterances, enabling the capture of complex interaction patterns beyond traditional pairwise graphs. A dual-stream gated attention network (DSGAN) is introduced to reduce information redundancy among nodes. Furthermore, we design a Transformer-based pairwise cross-modal fusion (PCMF) mechanism, which treats each unimodal feature as an anchor in turn and performs pairwise fusion with other modalities to extract deep emotional interaction information. Experiments on standard multimodal datasets show that the framework maintains high emotion recognition accuracy while effectively protecting user privacy. It outperforms generic models in personalized scenarios, achieving better alignment with individual emotional expression habits. The proposed HLPCMF model effectively captures multimodal emotional cues and enhances fine-grained emotion recognition in conversational contexts. The integration of hypergraph learning and pairwise cross-modal fusion provides a robust framework for modeling complex dependencies in multimodal dialogue emotion recognition tasks.