MDLFormer: multi-modal global context-based swin transformer for image manipulation detection and localization
摘要
Robust image manipulation software proliferation has significantly facilitated digital image tampering, resulting in numerous security concerns. Hence, localizing the manipulated regions and identifying manipulated images is critical. Even with the current focus on image tampering localization, tampering localization remains challenging in real-world forensic applications. Most current methods utilize deep neural networks (DNNs) for image manipulation detection and localization, resulting in optimistic results. However, DNN-based methods have drawbacks when establishing long-range dependencies. Consequently, solving the image manipulation localization issue requires a solution that can effectively construct a global context while maintaining a strong understanding of fine-grained features. To tackle this issue, we present MDLFormer in this work. MDLFormer consists of multi-modal volumetric input data and a Global Context-based Swin Transformer that can effectively capture global dependencies and low-level artifacts. Furthermore, to obtain discriminative representations of tampering evidence, we employed an FPN-based decoder. Extensive experiments conducted on various standard datasets, namely Columbia, Coverage, CASIA, NIST16, and IMD20, demonstrate that our model performs better than current state-of-the-art (SOTA) image manipulation detection and localization methods.