Rumor spreaders are increasingly using multimedia content on social platforms to attract and mislead the public. To alleviate this issue, research on rumor detection has received extensive attention in recent years. Although current multimodal rumor detection models utilize both textual and visual features of rumors for detection, they overlook the problem of text-image mismatch in rumors and the differences in the extraction processes between textual data and image data (such as temporal information and spatial structure). This paper proposes a new rumor detection method based on Multi-modal Mixture of Experts with Cross-modal Enhancement (MMoE). This method uses a contrastive language-image pre-trained model to calculate the similarity between text and images. Through a gating unit, different experts are selected based on the similarity to ensure that the model's predictions are not misled by incorrect information. In addition, a visual encoder-decoder is employed to convert post images into descriptive text for data augmentation, leveraging additional text data to optimize the model's performance. Experimental results show that the rumor detection results of MMoE outperform six baselines and achieve an accuracy rate of 92.49% on a Chinese dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Rumor Detection Method Based on Multi-modal Mixture of Experts and Cross-modal Enhancement

  • Jianyong Yu,
  • Xiuyu Li,
  • Xue Han

摘要

Rumor spreaders are increasingly using multimedia content on social platforms to attract and mislead the public. To alleviate this issue, research on rumor detection has received extensive attention in recent years. Although current multimodal rumor detection models utilize both textual and visual features of rumors for detection, they overlook the problem of text-image mismatch in rumors and the differences in the extraction processes between textual data and image data (such as temporal information and spatial structure). This paper proposes a new rumor detection method based on Multi-modal Mixture of Experts with Cross-modal Enhancement (MMoE). This method uses a contrastive language-image pre-trained model to calculate the similarity between text and images. Through a gating unit, different experts are selected based on the similarity to ensure that the model's predictions are not misled by incorrect information. In addition, a visual encoder-decoder is employed to convert post images into descriptive text for data augmentation, leveraging additional text data to optimize the model's performance. Experimental results show that the rumor detection results of MMoE outperform six baselines and achieve an accuracy rate of 92.49% on a Chinese dataset.