NRAFN: a non-text reinforcement and adaptive fusion network for multimodal sentiment analysis
摘要
With the rise of short videos, there has been an increasing interest in multimodal sentiment analysis, which involves text, audio, and visual modalities. One of the main challenges in multimodal sentiment analysis is designing an efficient fusion network. Most of the previous studies often overlook the difference in information density among different modalities, the dominant modality varies across different video clips, and some original information may be lost in the fusion process. To address the aforementioned issues, this paper proposes a non-text reinforcement and adaptive fusion network(NRAFN). The model reinforces audio and visual modalities through cross-modal attention and then completes multimodal fusion using an adaptive fusion module. Simultaneously, we generate unimodal sentiment labels for multi-task learning through a self-supervised approach to improve the generalization of the model. We conduct extensive experiments on the CMU-MOSI and CMU-MOSEI datasets. The experimental results demonstrate that NRAFN achieves state-of-the-art performance. To the best of our knowledge, this is the first paper that combines adaptive fusion and multi-task learning in multimodal sentiment analysis.