Multimodal sentiment analysis (MSA) seeks to interpret human emotions by integrating text, images, and audio. Despite progress, existing methods struggle with class imbalance and ineffective cross-modal sentiment representation. We propose EMSA, an ensemble-based framework combining a primary and auxiliary model, both enhanced with image captioning and consistency loss. Caption generation via multimodal large language models strengthens visual-textual alignment, while consistency loss improves modality-specific sentiment learning. By training on both balanced and imbalanced datasets, EMSA mitigates class imbalance without distorting data distribution. Experiments on public benchmarks show that EMSA achieves state-of-the-art results, especially under class imbalance and complex cross-modal conditions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EMSA: An Ensemble-Based Framework for Multimodal Sentiment Analysis

  • Ning Wang,
  • Fangyu Wu,
  • Shan Liang,
  • Yuhao Zhu,
  • Chaoyi Pang

摘要

Multimodal sentiment analysis (MSA) seeks to interpret human emotions by integrating text, images, and audio. Despite progress, existing methods struggle with class imbalance and ineffective cross-modal sentiment representation. We propose EMSA, an ensemble-based framework combining a primary and auxiliary model, both enhanced with image captioning and consistency loss. Caption generation via multimodal large language models strengthens visual-textual alignment, while consistency loss improves modality-specific sentiment learning. By training on both balanced and imbalanced datasets, EMSA mitigates class imbalance without distorting data distribution. Experiments on public benchmarks show that EMSA achieves state-of-the-art results, especially under class imbalance and complex cross-modal conditions.