<p>Multimodal sentiment analysis has achieved notable progress in modeling cross-modal emotional dependencies. However, existing approaches still suffer from limitations in modality uncertainty modeling, fine-grained interaction, and cross-modal alignment. To address these challenges, this paper proposes a tri-modal expert alignment framework based on capsule-gated optimization, termed COTEA. Specifically, COTEA employs a variational capsule-driven sparse expert mechanism to dynamically model the uncertainty of the visual modality and adaptively regulate the fusion pathway, thereby effectively suppressing the influence of noisy features. In addition, a tri-modal cross-attention encoder (Tri-modal CME) is introduced to enable deep and symmetric interactions among text, audio, and visual modalities. Furthermore, Gromov–Wasserstein distance combined with Fréchet and Mutual Information Regularization (FAMIR) is incorporated to construct a multi-level cross-modal alignment mechanism, which promotes cross-modal representation consistency at both the local structural and global distributional levels. Experimental results on four public multimodal sentiment analysis benchmark datasets, including CMU-MOSI, CMU-MOSEI, CH-SIMS, and CH-SIMS v2, demonstrate that the proposed COTEA framework achieves competitive performance across multiple evaluation metrics. Further ablation studies, missing-modality analyses, and cross-dataset evaluations verify the effectiveness of the key components.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

COTEA: capsule-optimized tri-modal expert alignment for multimodal sentiment analysis

  • Zhiwen Liao,
  • Ke Qi,
  • Feixiang Yuan,
  • Jiaxiong Liu

摘要

Multimodal sentiment analysis has achieved notable progress in modeling cross-modal emotional dependencies. However, existing approaches still suffer from limitations in modality uncertainty modeling, fine-grained interaction, and cross-modal alignment. To address these challenges, this paper proposes a tri-modal expert alignment framework based on capsule-gated optimization, termed COTEA. Specifically, COTEA employs a variational capsule-driven sparse expert mechanism to dynamically model the uncertainty of the visual modality and adaptively regulate the fusion pathway, thereby effectively suppressing the influence of noisy features. In addition, a tri-modal cross-attention encoder (Tri-modal CME) is introduced to enable deep and symmetric interactions among text, audio, and visual modalities. Furthermore, Gromov–Wasserstein distance combined with Fréchet and Mutual Information Regularization (FAMIR) is incorporated to construct a multi-level cross-modal alignment mechanism, which promotes cross-modal representation consistency at both the local structural and global distributional levels. Experimental results on four public multimodal sentiment analysis benchmark datasets, including CMU-MOSI, CMU-MOSEI, CH-SIMS, and CH-SIMS v2, demonstrate that the proposed COTEA framework achieves competitive performance across multiple evaluation metrics. Further ablation studies, missing-modality analyses, and cross-dataset evaluations verify the effectiveness of the key components.