Integrating Contrastive Learning and Multimodal Fusion Network for Fake News Detection
摘要
In recent years, the rapid spread of fake news with textual and visual content on the Internet, While news containing both text and images tends to attract more attention than news that consists of text alone, which may mislead readers and have a negative impact. The attention of numerous researchers has been drawn to multimodal detection of fake news. Current models within this field consider not just textual and visual features but social structure features as well, resulting in promising performance in the recognition of fake news. Although the existing multimodal feature fusion methods consider the pairwise interactions between different modalities including text, image, and social graph, they ignore the high-order mutual coupling relationship of these three modalities, resulting in the lack of comprehensive inter-modal correlation capture in the fused multimodal representation, and cannot fully ensure the high coordination of different modalities at the semantic level. This study introduce a new multimodal feature-enhanced attention network utilizing contrastive learning for fake news detection, which considers the interactions among all three modalities simultaneously. In addition, we employ contrastive self-supervised learning to deeply explore the relationship between news with different labels, enrich the obtained representations by minimizing the distance in the semantic space between news samples and samples that share the same label, and enhance the cooperative consistency between modes. Our model demonstrates superior performance in detecting fake news when compared to existing techniques, as shown by the experimental results.