MICTE: Mutual Information and Cross-Modal Text Enhancement for Multimodal Sentiment Analysis
摘要
The main challenge of multimodal sentiment analysis (MSA) is to effectively integrate and optimize information from diverse modalities, such as textual, visual, and acoustic. This integration is crucial for achieving more accurate analysis and comprehension of human emotional states. However, the majority of previous studies have not fully extracted and utilized the salient information in the input data, nor have they thoroughly explored the intricate connections between different modalities, thus failing to accurately identify human sentiment. To tackle this challenge, we propose a framework named Mutual Information and Cross-modal Text Enhancement (MICTE) to resolve the modal fusion problem. First, we propose a cross-modal text-enhanced attention mechanism, which dynamically weights emotional information from text modality and propagates it to acoustic and visual modalities. This mechanism enhances the emotional expression capabilities of non-textual modalities by aligning their features with text-based emotional cues. Second, we utilize the mutual information maximization strategy to strengthen the intermodal associations at the input layer and to refine key task-related information at the fusion layer. This strategy eliminates redundant information and optimizes the representation of multimodal features. Consequently, it thoroughly explores the correlations between different modalities and significantly boosts the effectiveness of modal integration. The experimental results on the CMU-MOSI and CMU-MOSEI datasets validate the effectiveness and practicality of our model in handling complex sentiment analysis challenges.