Multimodal fake news has caused significant harm to economic and political systems, thereby emerging as a serious societal issue. Current methodologies predominantly concentrate on modeling superficial image features (e.g., texture and color characteristics). Although certain investigations have attempted to enhance detection performance by extracting deep semantic information through identification of entity objects within images, existing approaches still demonstrate insufficient representation learning of spatial positional relationships between entities. This limitation consequently restricts their capacity to effectively characterize complex semantic associations. Therefore, we propose a scene graph-based semantic enhancement framework (SGSE), which extracts relationships between entities to help the model deeply understand semantic information in images. Utilizing fine-grained information obtained through scene graphs, we employ a fusion mechanism based on text filtering to achieve relation-level alignment between images and text. Furthermore, we employ a cross-granularity integration module, which contains coarse-grained extraction and multi-granularity fusion, to ensure a comprehensive representation of fake news. Experiments conducted on three benchmark datasets consistently demonstrate the superiority of SGSE, achieving state-of-the-art performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Scene Graph-Based Semantic Enhancement for Multimodal Fake News Detection

  • Hongyun Ding,
  • Shuohao Li,
  • Hang Du,
  • Jiaxin Yang,
  • Zhong Yang,
  • Jun Zhang

摘要

Multimodal fake news has caused significant harm to economic and political systems, thereby emerging as a serious societal issue. Current methodologies predominantly concentrate on modeling superficial image features (e.g., texture and color characteristics). Although certain investigations have attempted to enhance detection performance by extracting deep semantic information through identification of entity objects within images, existing approaches still demonstrate insufficient representation learning of spatial positional relationships between entities. This limitation consequently restricts their capacity to effectively characterize complex semantic associations. Therefore, we propose a scene graph-based semantic enhancement framework (SGSE), which extracts relationships between entities to help the model deeply understand semantic information in images. Utilizing fine-grained information obtained through scene graphs, we employ a fusion mechanism based on text filtering to achieve relation-level alignment between images and text. Furthermore, we employ a cross-granularity integration module, which contains coarse-grained extraction and multi-granularity fusion, to ensure a comprehensive representation of fake news. Experiments conducted on three benchmark datasets consistently demonstrate the superiority of SGSE, achieving state-of-the-art performance.