A noise and edge extraction-based dual-branch method for shallowfake and deepfake localization
摘要
The trustworthiness of multimedia is being increasingly evaluated by advanced image manipulation localization (IML) techniques, resulting in the emergence of the IML field. To use artifacts, an effective manipulation model needs to separate manipulated and unmanipulated sections by non-semantic features. This requires direct comparisons between the two regions. Current models employ either feature approaches based on handcrafted features, convolutional neural networks (CNNs), or a hybrid approach that combines both. Handcrafted feature approaches assume that tampering has already happened, which makes them less helpful in dealing with different types of tampering. On the other hand, CNNs only capture semantic information, which is insufficient to deal with manipulation artifacts. In order to address these constraints, we developed a dual-branch model that integrates manually designed feature noise with conventional CNN features. The model employs a dual-branch strategy, where one branch integrates noise characteristics, and the other integrates RGB features using the hierarchical ConvNext Module. In addition, the model utilizes edge supervision loss to acquire boundary manipulation information, resulting in accurate localization at the edges. Furthermore, this architecture uses a feature augmentation module to optimize and refine the presentation of attributes. Extensive experimentation is done on the shallowfake datasets (CASIA, COVERAGE, COLUMBIA, NIST16) and deepfake faceforensics dataset to demonstrate their outstanding ability to extract features and superior performance compared to other baseline models. The model even reached the AUC score of 99%. Results show that the model is superior in comparison and easily outperforms the existing state-of-the-art (SoTA) models.