Hierarchical prompt fusion and image denoising for multimodal aspect-based sentiment analysis
摘要
Multimodal Aspect-Based Sentiment Analysis (MABSA) is a fine-grained task that aims to analyze users’ sentiment polarity towards target aspects through different and rich modal contents. In conjunction with this topic, many methods have been proposed to link modalities to form interactive judgments of emotional tendencies. However, the currently proposed methods have some disadvantages: (1) Since some image modalities are unrelated to text modalities, information irrelevant to the aspect will be introduced; (2) Different modalities are difficult to complement each other during the fusion process, resulting in poor fusion performance, and the fusion process may also introduce additional noise. To address these issues, we propose a novel MABSA network model that combines a simple noise filtering approach with an innovative fast learning method for effective classification. Specifically, for the visual modality, we introduce an image content filtering (ICF) layer to filter out information that is not relevant to the aspect and text modalities. In addition, in the fusion stage, we proposed an excellent bidirectional interactive prompt fusion (BIPF) layer, which integrates prompt learning into the query and fusion stages, realizes the fusion of different modalities and contextual semantic information, and makes modal fusion more comprehensive. The outcomes of our experiments indicate that the model we developed attains leading-edge performance levels when tested on two MABSA datasets. Furthermore, a large number of experiments have shown that the model we put forward exhibits remarkable performance and robustness.