Explanation-Inspired Transferable Adversarial Attacks with Layer-Wise Increment Decomposition
摘要
Adversarial attacks have gained significant attention in the context of neural network security. In the realm of black-box attacks, feature-level attack methods have substantially enhanced the transferability of adversarial examples. Nevertheless, existing approaches for assessing feature importance exhibit certain weaknesses. In this paper, an explanation method termed Layer-wise Increment Decomposition (LID) for calculating neuron relevance is firstly revisited by further combining it with SoftMax Gradient-LRP (SG-LRP) and Integrated Gradients (IG), making it a more robust and precise tool for guiding adversarial attacks. Building upon this foundation and drawing inspiration from existing intermediate-layer attacks, an alternative transferable attacking loss is proposed by naively adapting the LID-based neuron relevance for intermediate layers. By further incorporating sophisticated numerical schemes for the LID-induced loss, we enhance the transferability of adversarial examples. A series of experiments conducted on both normal and defense models demonstrate that the proposed approach either outperforms or achieves comparable transferability to state-of-the-art methods.