Multimodal graph contrastive learning for fake news video detection
摘要
As online video news gains more popularity for information acquisition, the presence of fake news video poses a new threat to users seeking the real news. Though previous studies have made much progress in detection fake news in text and image formats, video-formed fake news brings new and unique challenges: 1) the modality heterogeneity of fake news video, 2) the inherent data non-alignment of different modalities. However, existing methods simply fuse the multimodal feature at the decision layer or feature layer, and fail to explore the intrinsic complex correlations between different modalities. In this paper, we propose a multimodal graph contrastive learning framework which learning complex relations between different modalities for detecting fake news video. Specifically, to solve the issue of modality heterogeneity, we construct unimodal homogeneous graphs, cross-modal heterogeneous graphs which model the hidden relations within and across modalities, respectively. We design graph contrastive learning module to obtain graph representation without explicitly aligning the data which captures the underlying correlation across modalities. We conduct experiments on the public dataset and the results show that our proposed model outperforms existing methods on fake news video detection.