<p>Current research on multimodal conversation emotion analysis primarily focuses on modeling the context within a dialogue and employing attention-based multimodal fusion. However, it has not effectively explored cross-dialogue associations, which leads to an underutilization of complementary semantic information from similar dialogue scenes. This paper addresses these gaps by proposing the <b>C</b>r<b>O</b>ss-dialogue <b>S</b>cene <b>I</b>nteractive <b>K</b>nowledge <b>E</b>nhancement model (COSIKE), which enhances dialogue modeling by integrating cross-dialogue cross-modal scenes and commonsense knowledge information. In particular, (1) COSIKE constructs global and local scene interaction graphs based on enriched scene descriptions generated by large language models to explore inter- and intra-dialogue associations. (2) An overlapping graph-based multi-scene interaction learning is proposed for scene information transfer. (3) Cross-modal commonsense distillation is employed for knowledge enhancement. Extensive experiments on the MELD and M3ED datasets demonstrate that COSIKE outperforms state-of-the-art methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-dialogue scene interactive knowledge enhancement for multimodal conversation emotion analysis

  • Jian Liao,
  • Xiaoyu Wang,
  • Yu Feng,
  • Suge Wang,
  • Jianxing Zheng

摘要

Current research on multimodal conversation emotion analysis primarily focuses on modeling the context within a dialogue and employing attention-based multimodal fusion. However, it has not effectively explored cross-dialogue associations, which leads to an underutilization of complementary semantic information from similar dialogue scenes. This paper addresses these gaps by proposing the CrOss-dialogue Scene Interactive Knowledge Enhancement model (COSIKE), which enhances dialogue modeling by integrating cross-dialogue cross-modal scenes and commonsense knowledge information. In particular, (1) COSIKE constructs global and local scene interaction graphs based on enriched scene descriptions generated by large language models to explore inter- and intra-dialogue associations. (2) An overlapping graph-based multi-scene interaction learning is proposed for scene information transfer. (3) Cross-modal commonsense distillation is employed for knowledge enhancement. Extensive experiments on the MELD and M3ED datasets demonstrate that COSIKE outperforms state-of-the-art methods.