Multimodal Aspect-Based Sentiment Analysis (MABSA) is a fine-grained mission that intents to analyze users’ sentiment tendency towards target aspects through different and rich modal contents. In conjunction with this topic, many methods have been proposed to link modalities to form interactive judgments of emotional tendencies. However, the currently proposed methods have some disadvantages: (1) Since some image modalities are unrelated to text modalities, information irrelevant to the aspect will be introduced; (2) Different modalities are difficult to complement each other during the fusion process, resulting in poor fusion performance, and the fusion process may also introduce additional noise. To resolve these points, we propose a novel MABSA network model that combines a simple noise filtering approach with an innovative fast learning method for effective classification. Specifically, for the visual modality, we introduce an image content filtering (ICF) layer to filter out information that is not relevant to the aspect and text modalities. In addition, in the fusion stage, we proposed an excellent bidirectional interactive prompt fusion (BIPF) layer, which integrates prompt learning into the query and fusion stages, realizes the fusion of different modalities and contextual semantic information, and makes modal fusion more comprehensive. The outcomes of our experiments indicate that the model we developed attains leading-edge performance levels when tested on two MABSA datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bidirectional Interactive Prompt Fusion and Noise Filtering for Multimodal Aspect-Based Sentiment Analysis

  • Fuxian Zhu,
  • Xiaoli Xu

摘要

Multimodal Aspect-Based Sentiment Analysis (MABSA) is a fine-grained mission that intents to analyze users’ sentiment tendency towards target aspects through different and rich modal contents. In conjunction with this topic, many methods have been proposed to link modalities to form interactive judgments of emotional tendencies. However, the currently proposed methods have some disadvantages: (1) Since some image modalities are unrelated to text modalities, information irrelevant to the aspect will be introduced; (2) Different modalities are difficult to complement each other during the fusion process, resulting in poor fusion performance, and the fusion process may also introduce additional noise. To resolve these points, we propose a novel MABSA network model that combines a simple noise filtering approach with an innovative fast learning method for effective classification. Specifically, for the visual modality, we introduce an image content filtering (ICF) layer to filter out information that is not relevant to the aspect and text modalities. In addition, in the fusion stage, we proposed an excellent bidirectional interactive prompt fusion (BIPF) layer, which integrates prompt learning into the query and fusion stages, realizes the fusion of different modalities and contextual semantic information, and makes modal fusion more comprehensive. The outcomes of our experiments indicate that the model we developed attains leading-edge performance levels when tested on two MABSA datasets.