The abuse of image editing software has brought about a crisis of trust and potential security risks to society, making the design of a universal image manipulation localization (IML) method crucial in practical terms. Edge artifacts in tampered images serve as vital clues for localizing tampered regions, and the ability to learn these semantically agnostic features is pivotal. Swin Transformer combines the global modeling capabilities of transformer for long-range dependencies with the contextual feature-focusing abilities of Convolutional Neural Networks within local receptive fields, exhibiting strong feature representation capabilities. However, its non-hierarchical network structure makes it challenging to address multiscale features. To effectively focus on edge features and achieve tampered region localization and detection, we propose a multiscale edge-supervised vision transformer (MEViT) network for IML. MEViT achieves multiscale extraction of edge features by designing a feature pyramid network suitable for ViT. The extracted features are input into an edge supervision module, which uses attention mechanism to learn the distribution of tampered area edges. Extensive experiments on five benchmark datasets have demonstrated the superiority and generalization of our model. Moreover, it exhibits better robustness against JPEG compression and Gaussian blur.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MEViT : Multiscale Edge-Supervised Vision Transformer for Image Manipulation Localization

  • Ruijie Wu,
  • Wei Guo,
  • Yi Liu,
  • Hui Ren

摘要

The abuse of image editing software has brought about a crisis of trust and potential security risks to society, making the design of a universal image manipulation localization (IML) method crucial in practical terms. Edge artifacts in tampered images serve as vital clues for localizing tampered regions, and the ability to learn these semantically agnostic features is pivotal. Swin Transformer combines the global modeling capabilities of transformer for long-range dependencies with the contextual feature-focusing abilities of Convolutional Neural Networks within local receptive fields, exhibiting strong feature representation capabilities. However, its non-hierarchical network structure makes it challenging to address multiscale features. To effectively focus on edge features and achieve tampered region localization and detection, we propose a multiscale edge-supervised vision transformer (MEViT) network for IML. MEViT achieves multiscale extraction of edge features by designing a feature pyramid network suitable for ViT. The extracted features are input into an edge supervision module, which uses attention mechanism to learn the distribution of tampered area edges. Extensive experiments on five benchmark datasets have demonstrated the superiority and generalization of our model. Moreover, it exhibits better robustness against JPEG compression and Gaussian blur.