错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DRSS: a multimodal sentiment analysis approach based on dual representation and self-supervised learning strategy

  • Jing Meng,
  • Zhenfang Zhu,
  • Jiangtao Qi,
  • Huaxiang Zhang

摘要

Multimodal sentiment analysis (MSA) is crucial for constructing the complex relationship across modalities and learning effective representations of multimodal information. However, current methods primarily focus on consistent information across modalities, neglecting the specific information unique to each modality. Additionally, most MSA datasets lack annotations for individual modality, limiting the thorough exploration of modality-specific information. Therefore, for the aim of modeling multimodal representations and capturing modality-specific information effectively, we propose a novel approach for MSA based on dual representation and self-supervised strategy (DRSS). Firstly, each modality input is decomposed into modality-consistent and modality-specific representation using a modality-dual-representation (MDR) learning network, which employs a shared consistency learning network to capture modality-consistent features and distinct specificity learning networks to extract modality-specific features. By contrasting modality-consistent and -specific representations within the same sample, the model’s representational capacity is further enhanced. Furthermore, the unimodal label generation module (ULGM) based on self-supervised learning is employed to generate unimodal labels for each modality, preserving the differentiation between modalities through unimodal label prediction. Finally, a multitask prediction loss is introduced to jointly train multimodal, unimodal, and contrastive learning tasks. The optimization of modality consistency and specificity losses is integrated into a unified loss function. Additionally, a self-adjusting strategy is employed in unimodal tasks to dynamically update the loss weights, guiding the model to effectively learn modality specificity. Experimental results on the CMU-MOSI and CMU-MOSEI datasets, which lack unimodal annotations, demonstrate that DRSS achieves competitive performance across a range of metrics.