Multimodal Sentiment Analysis (MSA) is a popular research field in natural language processing and affective computing. Recent works have shown the feasibility of Multimodal Large Language Models (MLLMs) for zero-shot MSA. However, we observe that existing MLLMs-based work has the following limitations. (1) The responses do not contain the required labels. (2) There exists the bias brought by the order of label options in prompts. Moreover, for some ambiguous instances, it is difficult to answer correctly by directly asking MLLMs. Thus, we propose a novel MLLMs-based Two-stage Model (M2S) to help distinguish those hard samples. The first stage aims to select hard samples using our proposed probability distribution-based method. The second stage aims to predict sentiment polarity for hard samples by using auxiliary image content with our proposed ensembled method. We also design two new prompts with three kinds of label option orders for both fine-grained and coarse-grained MSA tasks. Experiments on five MSA datasets with different granularities show that our M2S method has substantial improvements by 6% accuracy on average over the state-of-the-art method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Novel MLLMs-Based Two-Stage Model for Zero-Shot Multimodal Sentiment Analysis

  • Heng-Yang Lu,
  • Xiao-Fei Li

摘要

Multimodal Sentiment Analysis (MSA) is a popular research field in natural language processing and affective computing. Recent works have shown the feasibility of Multimodal Large Language Models (MLLMs) for zero-shot MSA. However, we observe that existing MLLMs-based work has the following limitations. (1) The responses do not contain the required labels. (2) There exists the bias brought by the order of label options in prompts. Moreover, for some ambiguous instances, it is difficult to answer correctly by directly asking MLLMs. Thus, we propose a novel MLLMs-based Two-stage Model (M2S) to help distinguish those hard samples. The first stage aims to select hard samples using our proposed probability distribution-based method. The second stage aims to predict sentiment polarity for hard samples by using auxiliary image content with our proposed ensembled method. We also design two new prompts with three kinds of label option orders for both fine-grained and coarse-grained MSA tasks. Experiments on five MSA datasets with different granularities show that our M2S method has substantial improvements by 6% accuracy on average over the state-of-the-art method.