Towards robust multimodal emotion recognition in conversation with multi-modal transformer and variational distillation fusion
摘要
Multimodal Emotion Recognition in Conversation (MERC) utilizes multimodal information such as language, visual, and audio to enhance the understanding of human emotions. Current multimodal interaction frameworks inadequately resolve inherent information conflicts and redundancy due to their assumption of equivalent quality across heterogeneous modalities. In addition, inappropriate evaluation of the importance of modalities can also cause this problem. To address this issue, we introduce a