Nowadays, modern social networks allow the rapid sharing of news worldwide. Alas, these news are frequently unverified or shared on the basis of users’ opinions or beliefs, which can cause confusion widespread, public trust erosion, and contribution to social and political instability. In this complex and evolving scenario, the early detection of fake news has become a critical issue. Multimodal approaches, which integrate various data types such as text, images, audio, video, and network structures, have shown promising results in addressing such a problem. The literature presents different fusion strategies, but there is no consensus on which one is the most effective. In this work, we propose \(M3DUSA\) , a modular multi-modal framework able to combine different modalities to effectively detect malicious and misleading content. By using deep attention-based architectures, our framework discovers informative latent representations that can be combined using different early or late fusion strategies. Experiments conducted on a real-world dataset demonstrate the effectiveness of our solution. The achieved results highlight that while both early and late fusion approaches can effectively exploit the complementary contributions from different modalities, they can exhibit distinct advantages depending on the desired outcomes.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

You Can Spread But You Cannot Hide: Discovering Accurate Multi-modal Deep Fusion Models for Fake News Detection

  • Liliana Martirano,
  • Paolo Zicari,
  • Massimo Guarascio,
  • Francesco Sergio Pisani,
  • Carmela Comito

摘要

Nowadays, modern social networks allow the rapid sharing of news worldwide. Alas, these news are frequently unverified or shared on the basis of users’ opinions or beliefs, which can cause confusion widespread, public trust erosion, and contribution to social and political instability. In this complex and evolving scenario, the early detection of fake news has become a critical issue. Multimodal approaches, which integrate various data types such as text, images, audio, video, and network structures, have shown promising results in addressing such a problem. The literature presents different fusion strategies, but there is no consensus on which one is the most effective. In this work, we propose \(M3DUSA\) , a modular multi-modal framework able to combine different modalities to effectively detect malicious and misleading content. By using deep attention-based architectures, our framework discovers informative latent representations that can be combined using different early or late fusion strategies. Experiments conducted on a real-world dataset demonstrate the effectiveness of our solution. The achieved results highlight that while both early and late fusion approaches can effectively exploit the complementary contributions from different modalities, they can exhibit distinct advantages depending on the desired outcomes.