Multi-Dimensional Spatiotemporal Modeling for Multimodal Emotion Recognition in Conversations
摘要
Multimodal Emotion Recognition in Conversation (MERC) aims to analyze emotions by integrating audio, text, and video data. Accurately recognizing emotions requires modeling the contextual interdependence among utterances. However, most existing methods focus on pairwise utterance interactions, neglecting global context. Moreover, conversations unfold in a chronological order and are inherently composed of three modalities, collectively forming a unified spatial-temporal structure. Independently modeling temporal and spatial dimensions may violate the intrinsic interdependence between time and space in real-world interaction. To address these challenges, we propose a Multi-Dimensional Spatiotemporal Modeling (MSM) approach. First, we construct a high-order conversation interaction graph, leveraging simplex structures to extract multi-dimensional semantic information. A bipartite graph is further introduced to capture high-order interactions between the target utterance and one or more related utterances. To enhance information propagation, we employ a high-order graph convolutional network with dedicated high-frequency and low-frequency filters, projecting semantic signals from conversations into the frequency domain. This allows us to decompose emotion-related information across different frequency bands while maintaining a unified spatiotemporal representation. The adaptive filtering mechanism further integrates signals from different frequency ranges, improving the model’s capability to capture cross-modal and cross-utterance spatiotemporal interactions. Extensive experiments conducted on two widely used MERC datasets demonstrate that our proposed method outperforms the current state-of-the-art approaches.