<p>Deep learning has achieved remarkable results in the field of Image Forgery Localization (IFL), but the existing methods pay insufficient attention to the training efficiency and the spatial structure and size of the tampered region. To address this limitation, we propose the Multi-scale Query-based Transformer (MQFormer) model, that employs Ground-truth Mask Token (GMT) to facilitate the identification of forged regions using Image Feature Token (IFT) and Learnable Query Token (LQT). In particular, IFT are initially extracted through the utilization of a dual-stream encoder (comprising RGB and Noise branches), while the feature embedding of the ground-truth mask is then employed as GMT. The key contribution of MQFormer is the design of a multi-scale query transformer module, which leverage attention mechanism to update IFT and LQT through GMT. The module continuously optimizes the localization accuracy from coarse to fine grains through gradual refinement, thus enhancing the model’s ability to identify forged regions of different sizes. Furthermore, we propose a novel contrastive loss function, which aims to reduce the distance between GMT and LQT in the feature space, thereby enhancing the guiding role of mask information. Extensive experimentation on multiple benchmarks has demonstrated that MQFormer markedly outperforms existing state-of-the-art methodologies and displays adaptability to diverse tampering types and robust resilience to attacks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-scale query-based transformer for image forgery localization

  • Ruyi Bai

摘要

Deep learning has achieved remarkable results in the field of Image Forgery Localization (IFL), but the existing methods pay insufficient attention to the training efficiency and the spatial structure and size of the tampered region. To address this limitation, we propose the Multi-scale Query-based Transformer (MQFormer) model, that employs Ground-truth Mask Token (GMT) to facilitate the identification of forged regions using Image Feature Token (IFT) and Learnable Query Token (LQT). In particular, IFT are initially extracted through the utilization of a dual-stream encoder (comprising RGB and Noise branches), while the feature embedding of the ground-truth mask is then employed as GMT. The key contribution of MQFormer is the design of a multi-scale query transformer module, which leverage attention mechanism to update IFT and LQT through GMT. The module continuously optimizes the localization accuracy from coarse to fine grains through gradual refinement, thus enhancing the model’s ability to identify forged regions of different sizes. Furthermore, we propose a novel contrastive loss function, which aims to reduce the distance between GMT and LQT in the feature space, thereby enhancing the guiding role of mask information. Extensive experimentation on multiple benchmarks has demonstrated that MQFormer markedly outperforms existing state-of-the-art methodologies and displays adaptability to diverse tampering types and robust resilience to attacks.