Interpretable and robust multimodal fake news detection with evidence fusion and reliable reasoning
摘要
Multi-modal fake news detection involves the complex task of jointly analyzing textual and visual information, whereas uni-modal systems often under-perform due to limited contextual understanding. This paper presents a reliable and interpretable multi-modal framework that integrates semantic and stylistic cues using BERT for textual encoding, and CLIP and SWIN Transformer models for visual representation. Intra-modal features are fused via a Dempster–Shafer theory-based gating mechanism, enabling uncertainty-aware reasoning within each modality. Cross-modal features are aligned and refined using transformer-based correlation modeling to capture nuanced interactions between modalities. The framework is further strengthened by a composite loss function, formed by combining evidential deep learning with supervised and unsupervised contrastive learning strategies. This results in improving the classification accuracy and enhancing the model’s robustness by quantifying uncertainty and allowing the rejection of low-confidence predictions. The inclusion of the uncertainty regularization ensures well-calibrated outputs, which are crucial for real-world deployment where interpret-ability and trust are essential. Extensive experiments on the Fakeddit dataset demonstrate that the proposed approach consistently outperforms strong baselines in terms of accuracy, uncertainty estimation, and transparency of decisions. By tightly coupling multi-modal evidence integration with principled uncertainty modeling, the framework advances the state of the art in trustworthy misinformation detection.