With the continuous development of multimedia, Multimodal Sentiment Analysis has become a highly regarded field. Recent research proposes learning effective unimodal representations to facilitate multimodal fusion, which mainly contain two parts of information: modality-common and modality-specific information. However, previous work does not consider the sentiment span information between different samples during the fusion process. In this paper, we propose a novel framework Weighted Contrastive Fusion (WConF) to extract sentiment span information for multimodal fusion. First, we apply modality contrastive learning to capture modality-common sentiment information and separate modality-specific sentiment information. Then, considering sentiment spans as weights, we perform weighted contrastive learning on the multimodal representations. Moreover, we design a loss function to assist in weighted contrastive fusion. Finally, we conduct extensive experiments on four benchmark datasets: MOSI, MOSEI, CH-SIMS, and CH-SIMSV2. Experimental results demonstrate that our model achieves superior performance compared to previous models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

WConF: Weighted Contrastive Fusion for Multimodal Sentiment Analysis

  • Biqing Zeng,
  • Ruiyuan Li,
  • Liuxing Lu,
  • Liangqi Xie,
  • Jiazhen Wang,
  • Weihai Chen,
  • Huimin Deng

摘要

With the continuous development of multimedia, Multimodal Sentiment Analysis has become a highly regarded field. Recent research proposes learning effective unimodal representations to facilitate multimodal fusion, which mainly contain two parts of information: modality-common and modality-specific information. However, previous work does not consider the sentiment span information between different samples during the fusion process. In this paper, we propose a novel framework Weighted Contrastive Fusion (WConF) to extract sentiment span information for multimodal fusion. First, we apply modality contrastive learning to capture modality-common sentiment information and separate modality-specific sentiment information. Then, considering sentiment spans as weights, we perform weighted contrastive learning on the multimodal representations. Moreover, we design a loss function to assist in weighted contrastive fusion. Finally, we conduct extensive experiments on four benchmark datasets: MOSI, MOSEI, CH-SIMS, and CH-SIMSV2. Experimental results demonstrate that our model achieves superior performance compared to previous models.