Multilayer interactive attention bottleneck transformer for aspect-based multimodal sentiment analysis
摘要
Currently, aspect-based multimodal sentiment analysis remains a highly hot research field, aiming to leverage various modalities such as images and text to determine the sentiment orientation of viewpoint entities. Although deep learning methods have made significant progress in this field, some challenges still exist: incomplete alignment of information between modalities, insufficient interaction and low utilization during modal fusion. To solve these problems, this paper proposes a novel multimodal sentiment analysis model called multilayer interactive attention bottleneck transformer (MIABT) network model. The model contains two key modules: (1) the first is the Multimodal Dynamic Gate (MDG) module, which can dynamically interact to align image features and text features; (2) the second is the Multimodal Attention Bottleneck Transformer (MABT) module, which improves performance at lower computational costs by limiting the flow of information between modalities, only sharing necessary relevant information to restrict cross-modal attention. Experimental results show that the model outperforms the baseline model on two public datasets, Twitter-2015 and Twitter-2017, demonstrating that our proposed approach effectively enhances the accuracy of aspect-based multimodal sentiment analysis tasks.