Multimodal Sentiment Analysis (MSA) has gained attention in fields such as video understanding and risk management due to its ability to integrate multiple information modalities, including text, visual, and audio. However, existing MSA methods face challenges in balancing cross-modal consistency with modality-specific characteristics. Moreover, these methods often rely on manually annotated unimodal data, which incurs substantial costs and operational difficulties. In this study, we propose an improved self-supervised multitask learning framework, named Self-supervised Uni-Multi modal Sentiment Analysis Net (SUMSA Net), to address these challenges. Our approach incorporates a feature similarity-based Unimodal Pseudo-Label Generation (UPLG) module to automatically generate unimodal labels, thereby reducing dependency on manual annotations. Additionally, we introduce an Interaction Modal Fusion Module (IFM) to effectively capture complex nonlinear interactions between modalities. Furthermore, xLSTM is integrated into the framework to better capture temporal dependencies in sequential data from audio and video modalities. By jointly learning multimodal and unimodal tasks, SUMSA Net captures both shared and modality-specific features. Experimental results on the MOSI, MOSEI, and SIMS datasets demonstrate that SUMSA Net achieves notable performance improvements in sentiment analysis while significantly reducing the need for manual annotations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Multimodal Sentiment Analysis with Unimodal Pseudo-label Generation

  • Jianing Zhao,
  • Ou Deng,
  • Qun Jin

摘要

Multimodal Sentiment Analysis (MSA) has gained attention in fields such as video understanding and risk management due to its ability to integrate multiple information modalities, including text, visual, and audio. However, existing MSA methods face challenges in balancing cross-modal consistency with modality-specific characteristics. Moreover, these methods often rely on manually annotated unimodal data, which incurs substantial costs and operational difficulties. In this study, we propose an improved self-supervised multitask learning framework, named Self-supervised Uni-Multi modal Sentiment Analysis Net (SUMSA Net), to address these challenges. Our approach incorporates a feature similarity-based Unimodal Pseudo-Label Generation (UPLG) module to automatically generate unimodal labels, thereby reducing dependency on manual annotations. Additionally, we introduce an Interaction Modal Fusion Module (IFM) to effectively capture complex nonlinear interactions between modalities. Furthermore, xLSTM is integrated into the framework to better capture temporal dependencies in sequential data from audio and video modalities. By jointly learning multimodal and unimodal tasks, SUMSA Net captures both shared and modality-specific features. Experimental results on the MOSI, MOSEI, and SIMS datasets demonstrate that SUMSA Net achieves notable performance improvements in sentiment analysis while significantly reducing the need for manual annotations.