The fusion of multimodal cues, i.e., visual, audio, and language, can provide complementary insights and benefit sentiment analysis. However, not all of them are always available in practical scenarios, posing a challenge for inference with incomplete modality. One idea to address this issue is tailoring distillation algorithms to transfer multimodal knowledge from a full-modality teacher to the incomplete-modality student for missing information compensation. However, existing works utilize fixed and unified knowledge from the pretrained teacher to guide students under different missing states, ignoring their varying capacities. Consequently, the knowledge may be challenging for students with lower capacities to assimilate due to the huge gap. Thus, we propose a missing-customized distillation framework, to adapt the teacher’s knowledge to different missing-state students for better knowledge transfer. Specifically, for the student side, we devise a learnable missing token for each modality to perceive the current missing status. Then for the teacher side, these missing-aware tokens, along with extra trainable adapters, are inserted into it to facilitate knowledge adaptation. With a novel gradient-guided interactive training strategy, our method ensures that the teacher provides student-adaptive knowledge for different students and then the student absorbs valuable knowledge for improved missing information recovery. Extensive experiments on CMU-MOSI and CMU-MOSEI datasets validate the efficiency of our method compared with the state-of-the-arts across various missing rates.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Missing Customized Distillation Network for Incomplete Multimodal Sentiment Analysis

  • Zhangfeng Hu,
  • Wenming Zheng,
  • Mengting Wei,
  • Mengxin Shi,
  • Yuan Zong

摘要

The fusion of multimodal cues, i.e., visual, audio, and language, can provide complementary insights and benefit sentiment analysis. However, not all of them are always available in practical scenarios, posing a challenge for inference with incomplete modality. One idea to address this issue is tailoring distillation algorithms to transfer multimodal knowledge from a full-modality teacher to the incomplete-modality student for missing information compensation. However, existing works utilize fixed and unified knowledge from the pretrained teacher to guide students under different missing states, ignoring their varying capacities. Consequently, the knowledge may be challenging for students with lower capacities to assimilate due to the huge gap. Thus, we propose a missing-customized distillation framework, to adapt the teacher’s knowledge to different missing-state students for better knowledge transfer. Specifically, for the student side, we devise a learnable missing token for each modality to perceive the current missing status. Then for the teacher side, these missing-aware tokens, along with extra trainable adapters, are inserted into it to facilitate knowledge adaptation. With a novel gradient-guided interactive training strategy, our method ensures that the teacher provides student-adaptive knowledge for different students and then the student absorbs valuable knowledge for improved missing information recovery. Extensive experiments on CMU-MOSI and CMU-MOSEI datasets validate the efficiency of our method compared with the state-of-the-arts across various missing rates.