错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A cross-modal imagination network based on joint calibration for multimodal sentiment analysis with missing modalities

  • Chen Hong,
  • Xianyong Li,
  • Dong Huang,
  • Yajun Du,
  • Yanli Lee,
  • Jia Liu,
  • Xiaoliang Chen,
  • Yongquan Fan

摘要

Multimodal sentiment analysis with missing modalities aims to predict sentiment intensity from incomplete multimodal inputs. However, existing methods often model temporally unaligned cross-modal interaction insufficiently, and the recovered missing modality representations may remain shallow, limiting the preservation of deep sentiment semantics. To address these issues, this paper proposes a cross-modal imagination network based on joint calibration, which consists of the missing modality generation module, the cross-modal interaction learning module, and the temporal condition reconstruction module. Specifically, missing modality representations are first generated from the existing modalities. A triple-supervision calibration strategy is then introduced through modal knowledge distillation, representation similarity, and semantic alignment, so that modality-specific information can be preserved while modality-invariant sentiment cues can also be retained for subsequent fusion. The generated and existing representations are further refined by a shared semantic encoder and then fed into the cross-modal interaction learning module for deep interaction over unaligned sequences. In parallel, the temporal condition reconstruction module serves as an auxiliary branch to refine incomplete latent representations by jointly modeling local temporal continuity and global sequential dependency. Because robust multimodal learning under incomplete and temporally unaligned conditions involves repeated interaction modeling and joint optimization, the framework is also relevant to high-performance intelligent computing, where GPU-accelerated and parallel computation can improve large-scale training efficiency. Experiments on two public datasets show that CINJC improves the F1 score by 0.7−11.4% and 0.4−18.4% under single- and dual-modal missing scenarios, respectively, compared with representative methods.