Cross-Modal Memory Attention Network with Multi-view for Multimodal Rumor Detection
摘要
As social media platforms continue to evolve and expand, inadvertently fostering an environment ripe for the proliferation of rumors. To alleviate the potential negative impact of rumors on society, researchers propose diverse automatic rumors detection methods. Regrettably, most of existing detection methods incline to extract features from a single view, overlooking the wealth information encapsulated within alternative views. Consequently, the extracted features harbor an inherent deficiency in informational completeness, which impact the performance of the methods. In this paper, we propose a novel approach called multi-view multimodal cross-modal memory attention network (MVCMA) for rumor detection. Multi-view features of textual modality and visual modality are extracted in this model to generate the final representations. The cross-modal memory attention is proposed to integrate single-view feature of one modality with multi-view features of another modality, tackling the challenge of merging individual with multiple features effectively. Besides, we employ unimodal memory attention approach and cosine similarity to quantify inconsistencies among multimodal features, mitigating its detrimental effects on performance. The performance of our proposed model is assessed using widely recognized datasets, which demonstrate the MVCMA model can perform better than state-of-the-art related multimodal rumor detection methods.