Exploring Cross-Modal Inconsistency in Entities and Emotions for Multimodal Fake News Detection
摘要
The automatic detection of multimodal fake news has attracted significant attention recently. Numerous existing methods focus on the fusion of unimodal features to generate multimodal news representations. However, it is possible that these methods have not successfully acquired aligned modal information with sufficient accuracy and failed to effectively leverage the entity inconsistency present across modalities. Besides, there has been a lack of exploration regarding the emotional inconsistency across modalities. To address that, we propose CINEMA, a novel framework to explore cross-modal inconsistency in entities and emotions for multimodal fake news detection. We leverage the cross-modal contrastive learning objective to establish the alignment between the image and text modalities. An entity consistency learning module is developed to learn the cross-modality entity correlations. An emotional consistency learning module is implemented to effectively capture the emotional information within each modality. Finally, we evaluate the performance of CINEMA and conduct a comparative study using two extensively used datasets, Twitter and Weibo. The experimental results unequivocally demonstrate that our proposed CINEMA framework surpasses previous approaches by a substantial margin, establishing new state-of-the-art results on both datasets.