EMSA: An Ensemble-Based Framework for Multimodal Sentiment Analysis
摘要
Multimodal sentiment analysis (MSA) seeks to interpret human emotions by integrating text, images, and audio. Despite progress, existing methods struggle with class imbalance and ineffective cross-modal sentiment representation. We propose EMSA, an ensemble-based framework combining a primary and auxiliary model, both enhanced with image captioning and consistency loss. Caption generation via multimodal large language models strengthens visual-textual alignment, while consistency loss improves modality-specific sentiment learning. By training on both balanced and imbalanced datasets, EMSA mitigates class imbalance without distorting data distribution. Experiments on public benchmarks show that EMSA achieves state-of-the-art results, especially under class imbalance and complex cross-modal conditions.