DRLN: Disentangled Representation Learning Network for Multimodal Sentiment Analysis
摘要
Multimodal sentiment analysis (MSA) involves the analysis of human emotions by integrating emotional features from text, audio, and visual modalities. Previously, a considerable amount of research had commonly focused on fusing these three modalities equally, neglecting the importance of text modality and the influence of modal heterogeneity on fusion. To address this issue, we propose the Disentangled Representation Learning Network (DRLN) which decomposes representations from different modalities into separate sub-spaces centered around text modality to capture similarity and dissimilarity features among modalities. Meanwhile, we utilize contrastive learning loss and reconstruction loss to better learn representations within each sub-space. Finally, extensive experiments on two benchmark datasets demonstrate that our model outperforms state-of-the-art methods on various metrics.