MSMAE-Net: multi-semantic and multi-attention enhanced network for image forgery localization
摘要
The popularization of modern digital image technology has brought convenience to us, but it also poses many risks. The advancement of image editing software allows anyone to modify image content effortlessly. If these modified images are abused, they can severely impact societal safety and stability. To address these risks, we propose an end-to-end multi-scale collaborative enhancement image forgery localization network, termed MSMAE-Net. The method first employs a multi-branch feature extractor to initially capture both global and local information, with each branch optimized for different contextual information and features of the image. Next, to enhance the model’s ability to capture forged traces at various levels while maintaining the correlation between different spatial hierarchies, a boundary information aggregation module is designed. By constructing multiple branches with different receptive fields and collaborative learning with each other, a nested collaborative enhancement branch is designed to make the backbone network pay more attention to key features, so as to obtain better feature representation capability. Furthermore, to further fuse different semantic information, a novel strongly compatible semantic attention fusion module is proposed in this paper. Finally, an attention enhancement module is introduced in the localization of forgery regions to adaptively concentrate on key forged areas in the image and effectively capture the relationships between different pixels. In the experiments, we demonstrate that MSMAE-Net exhibits significant advantages in both localization accuracy and robustness in complex forgery scenarios when compared to other SOTA image forgery detection methods.