Multimodal sentiment analysis (MSA) involves the analysis of human emotions by integrating emotional features from text, audio, and visual modalities. Previously, a considerable amount of research had commonly focused on fusing these three modalities equally, neglecting the importance of text modality and the influence of modal heterogeneity on fusion. To address this issue, we propose the Disentangled Representation Learning Network (DRLN) which decomposes representations from different modalities into separate sub-spaces centered around text modality to capture similarity and dissimilarity features among modalities. Meanwhile, we utilize contrastive learning loss and reconstruction loss to better learn representations within each sub-space. Finally, extensive experiments on two benchmark datasets demonstrate that our model outperforms state-of-the-art methods on various metrics.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DRLN: Disentangled Representation Learning Network for Multimodal Sentiment Analysis

  • Jingming Hou,
  • Nazlia Omar,
  • Sabrina Tiun,
  • Saidah Saad,
  • Qian He

摘要

Multimodal sentiment analysis (MSA) involves the analysis of human emotions by integrating emotional features from text, audio, and visual modalities. Previously, a considerable amount of research had commonly focused on fusing these three modalities equally, neglecting the importance of text modality and the influence of modal heterogeneity on fusion. To address this issue, we propose the Disentangled Representation Learning Network (DRLN) which decomposes representations from different modalities into separate sub-spaces centered around text modality to capture similarity and dissimilarity features among modalities. Meanwhile, we utilize contrastive learning loss and reconstruction loss to better learn representations within each sub-space. Finally, extensive experiments on two benchmark datasets demonstrate that our model outperforms state-of-the-art methods on various metrics.