Multi-view Adaptive Fusion Model for Multimodal Fake News Detection
摘要
In recent years, automated detection of multimodal fake news has gained a wide range of attention. Most of the existing research focuses only on the process of semantic information fusion between multimodal features, ignoring the interactions of intra-modal and inter-modal semantics as well as the impact of image physical information (e.g., image recompression and tampering) on fake news detection. How to effectively extract intra-modal and inter-modal high-level semantic information while capturing physical information from multimodal data and adaptively aggregating both remains a challenging problem. To solve this problem, we propose a multi-view adaptive fusion model (MAFM). Specifically, our model integrates semantic and physical information. To extract higher-order semantic information we synthesize two views: intra-modal semantic similarity and inter-modal semantic similarity. At the same time intra-modal semantic information is used to learn inter-modal higher-order fusion semantic information. Similarly, we integrate the two views, local and global, to mine physical information from frequency domain features. For both types of information, we use the information aggregation module for adaptive fusion for categorization. Extensive experiments on two publicly available datasets show that our method outperforms related multimodal fake news detection methods.