Multimodal sentiment analysis based on temporal perception and cross-modal interaction
摘要
With the increasing prevalence of multimodal content on social platforms, sentiment expression has become more nuanced and complex, posing challenges for traditional unimodal sentiment analysis approaches. To address this issue, a multimodal sentiment analysis framework is proposed, which integrates temporal modeling and cross-modal interaction strategies. The architecture employs a Temporal Convolutional Network (TCN) to capture both short- and long-range dependencies within each modality, ensuring the preservation of sequential information. A Bidirectional Cross-Attention module is then introduced to facilitate fine-grained alignment and interaction across textual, visual, and acoustic modalities. Additionally, a Contrastive Cross-Transformer component is incorporated to improve the quality of multimodal feature representations through contrastive learning. Experimental results on two widely used benchmark datasets, MOSI and MOSEI, demonstrate that the proposed method achieves moderate improvements in sentiment classification accuracy and F1-score when compared with competitive Transformer-based baselines. These findings indicate that combining temporal structure with cross-modal reasoning can effectively enhance the model’s performance in analyzing complex emotional expressions.