<p>Multimodal Sentiment Analysis (MSA) aims to integrate multiple information sources, such as text and visual modalities, to predict the sentiments expressed in the data. However, the scarcity of multimodal datasets poses a significant challenge for effective vision-language fusion. Prompt learning has recently emerged as a promising solution to this issue. Existing prompting methods either convert image regions into visual tokens that align with textual word dimensions using a visual encoder or directly use image features. However, these methods overlook the inherent differences between the characteristics of a visual encoder and a language model, making it challenging for the language model to directly comprehend the semantic information of the images. To address this, we propose an adaptive multimodal prompt-tuning approach for sentiment analysis. Specifically, in the proposed adaptive multimodal prompt generation module, we extract initial multimodal features using an advanced pre-trained model, ensuring comprehensive integration of information from different modalities. We then introduce and fuse a learnable vector with the initial multimodal features dynamically, creating contextually relevant multimodal prompts. Finally, we integrate these multimodal prompts with the maked text sequence vector and send them to a pre-trained language model to obtain the word probability distribution. This process enhances the language model’s ability to comprehend and utilize semantic information from multimodal inputs, especially images. Extensive experiments and analyses on two aspect-level and two sentence-level datasets demonstrate that our method outperforms existing state-of-the-art approaches, confirming the effectiveness of the proposed adaptive multimodal prompts.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive multimodal prompt-tuning model for few-shot multimodal sentiment analysis

  • Yan Xiang,
  • Anlan Zhang,
  • Junjun Guo,
  • Yuxin Huang

摘要

Multimodal Sentiment Analysis (MSA) aims to integrate multiple information sources, such as text and visual modalities, to predict the sentiments expressed in the data. However, the scarcity of multimodal datasets poses a significant challenge for effective vision-language fusion. Prompt learning has recently emerged as a promising solution to this issue. Existing prompting methods either convert image regions into visual tokens that align with textual word dimensions using a visual encoder or directly use image features. However, these methods overlook the inherent differences between the characteristics of a visual encoder and a language model, making it challenging for the language model to directly comprehend the semantic information of the images. To address this, we propose an adaptive multimodal prompt-tuning approach for sentiment analysis. Specifically, in the proposed adaptive multimodal prompt generation module, we extract initial multimodal features using an advanced pre-trained model, ensuring comprehensive integration of information from different modalities. We then introduce and fuse a learnable vector with the initial multimodal features dynamically, creating contextually relevant multimodal prompts. Finally, we integrate these multimodal prompts with the maked text sequence vector and send them to a pre-trained language model to obtain the word probability distribution. This process enhances the language model’s ability to comprehend and utilize semantic information from multimodal inputs, especially images. Extensive experiments and analyses on two aspect-level and two sentence-level datasets demonstrate that our method outperforms existing state-of-the-art approaches, confirming the effectiveness of the proposed adaptive multimodal prompts.