Multimodal Emotion Recognition in Conversation (MERC) aims to analyze emotions by integrating audio, text, and video data. Accurately recognizing emotions requires modeling the contextual interdependence among utterances. However, most existing methods focus on pairwise utterance interactions, neglecting global context. Moreover, conversations unfold in a chronological order and are inherently composed of three modalities, collectively forming a unified spatial-temporal structure. Independently modeling temporal and spatial dimensions may violate the intrinsic interdependence between time and space in real-world interaction. To address these challenges, we propose a Multi-Dimensional Spatiotemporal Modeling (MSM) approach. First, we construct a high-order conversation interaction graph, leveraging simplex structures to extract multi-dimensional semantic information. A bipartite graph is further introduced to capture high-order interactions between the target utterance and one or more related utterances. To enhance information propagation, we employ a high-order graph convolutional network with dedicated high-frequency and low-frequency filters, projecting semantic signals from conversations into the frequency domain. This allows us to decompose emotion-related information across different frequency bands while maintaining a unified spatiotemporal representation. The adaptive filtering mechanism further integrates signals from different frequency ranges, improving the model’s capability to capture cross-modal and cross-utterance spatiotemporal interactions. Extensive experiments conducted on two widely used MERC datasets demonstrate that our proposed method outperforms the current state-of-the-art approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-Dimensional Spatiotemporal Modeling for Multimodal Emotion Recognition in Conversations

  • Xiaoyang Wang,
  • Zhenyu Yang,
  • Xueli Chang,
  • Ziyu Chen,
  • Haozhi Xia

摘要

Multimodal Emotion Recognition in Conversation (MERC) aims to analyze emotions by integrating audio, text, and video data. Accurately recognizing emotions requires modeling the contextual interdependence among utterances. However, most existing methods focus on pairwise utterance interactions, neglecting global context. Moreover, conversations unfold in a chronological order and are inherently composed of three modalities, collectively forming a unified spatial-temporal structure. Independently modeling temporal and spatial dimensions may violate the intrinsic interdependence between time and space in real-world interaction. To address these challenges, we propose a Multi-Dimensional Spatiotemporal Modeling (MSM) approach. First, we construct a high-order conversation interaction graph, leveraging simplex structures to extract multi-dimensional semantic information. A bipartite graph is further introduced to capture high-order interactions between the target utterance and one or more related utterances. To enhance information propagation, we employ a high-order graph convolutional network with dedicated high-frequency and low-frequency filters, projecting semantic signals from conversations into the frequency domain. This allows us to decompose emotion-related information across different frequency bands while maintaining a unified spatiotemporal representation. The adaptive filtering mechanism further integrates signals from different frequency ranges, improving the model’s capability to capture cross-modal and cross-utterance spatiotemporal interactions. Extensive experiments conducted on two widely used MERC datasets demonstrate that our proposed method outperforms the current state-of-the-art approaches.