COTEA: capsule-optimized tri-modal expert alignment for multimodal sentiment analysis
摘要
Multimodal sentiment analysis has achieved notable progress in modeling cross-modal emotional dependencies. However, existing approaches still suffer from limitations in modality uncertainty modeling, fine-grained interaction, and cross-modal alignment. To address these challenges, this paper proposes a tri-modal expert alignment framework based on capsule-gated optimization, termed COTEA. Specifically, COTEA employs a variational capsule-driven sparse expert mechanism to dynamically model the uncertainty of the visual modality and adaptively regulate the fusion pathway, thereby effectively suppressing the influence of noisy features. In addition, a tri-modal cross-attention encoder (Tri-modal CME) is introduced to enable deep and symmetric interactions among text, audio, and visual modalities. Furthermore, Gromov–Wasserstein distance combined with Fréchet and Mutual Information Regularization (FAMIR) is incorporated to construct a multi-level cross-modal alignment mechanism, which promotes cross-modal representation consistency at both the local structural and global distributional levels. Experimental results on four public multimodal sentiment analysis benchmark datasets, including CMU-MOSI, CMU-MOSEI, CH-SIMS, and CH-SIMS v2, demonstrate that the proposed COTEA framework achieves competitive performance across multiple evaluation metrics. Further ablation studies, missing-modality analyses, and cross-dataset evaluations verify the effectiveness of the key components.