Emotion recognition plays a crucial role in various applications. Currently, most research primarily focuses on emotion recognition based on a single modality, such as speech, while neglecting information from other modalities like text and facial expressions. Moreover, existing methods typically focus only on either global or local information, whereas combining both can significantly improve accuracy. This paper proposes a Self-Distillation Model for Emotion Recognition with Cross-Modal Interaction and Graph Neural Networks (SDER). The model constructs local information through the Speaker Temporal GNN and captures global information via the Inter-modal and Intra-modal Interaction module. Additionally, based on the Transformer architecture, the model employs a hierarchical gating fusion mechanism to dynamically optimize the weight allocation between modalities, further enhancing emotion recognition capability. To strengthen the expressiveness of both local and global information, the model also introduces self-distillation learning, transferring knowledge from hard labels and soft labels to different modules. Experimental results on the IEMOCAP and CMU-MOSEI datasets demonstrate the effectiveness of the proposed method, significantly improving the accuracy of emotion recognition.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Self-Distillation Model for Emotion Recognition with Cross-Modal Interaction and Graph Neural Networks

  • Xiaomei Chen,
  • Jinfeng Wang,
  • Zhaozun Ou

摘要

Emotion recognition plays a crucial role in various applications. Currently, most research primarily focuses on emotion recognition based on a single modality, such as speech, while neglecting information from other modalities like text and facial expressions. Moreover, existing methods typically focus only on either global or local information, whereas combining both can significantly improve accuracy. This paper proposes a Self-Distillation Model for Emotion Recognition with Cross-Modal Interaction and Graph Neural Networks (SDER). The model constructs local information through the Speaker Temporal GNN and captures global information via the Inter-modal and Intra-modal Interaction module. Additionally, based on the Transformer architecture, the model employs a hierarchical gating fusion mechanism to dynamically optimize the weight allocation between modalities, further enhancing emotion recognition capability. To strengthen the expressiveness of both local and global information, the model also introduces self-distillation learning, transferring knowledge from hard labels and soft labels to different modules. Experimental results on the IEMOCAP and CMU-MOSEI datasets demonstrate the effectiveness of the proposed method, significantly improving the accuracy of emotion recognition.