Adversarial attacks manipulate Deep Neural Networks(DNNs) to produce inaccurate predictions by injecting imperceptible noise into input samples, posing a significant security risk to deep learning systems. While current adversarial attack techniques are proficient at assessing DNNs robustness with known model parameters, their ability to generalize to unfamiliar models is constrained, suggesting inadequate transferability of adversarial samples. This study proposes improving adversarial transferability via prediction feature, named PFA. Although DNNs have differences in parameters and structures, the prediction features generated for the same input samples are highly similar. PFA obtains and modifies these features in input samples to prompt misclassifications across multiple DNNs. First, the attention mechanism extracts discriminative prediction features from the DNNs. Subsequently, these prediction feature are structured into a loss function aimed at minimizing them. Through iterative optimization of the loss function, adversarial perturbations are generated. Finally, adversarial perturbations are incorporated into the input samples to create adversarial samples, enabling effective attacks on various DNNs. Extensive experiments conducted on ImageNet demonstrate that PFA achieves better transferability than typical adversarial attack methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PFA: Improving Adversarial Transferability via Prediction Feature

  • Pengju Wang,
  • Jing Liu

摘要

Adversarial attacks manipulate Deep Neural Networks(DNNs) to produce inaccurate predictions by injecting imperceptible noise into input samples, posing a significant security risk to deep learning systems. While current adversarial attack techniques are proficient at assessing DNNs robustness with known model parameters, their ability to generalize to unfamiliar models is constrained, suggesting inadequate transferability of adversarial samples. This study proposes improving adversarial transferability via prediction feature, named PFA. Although DNNs have differences in parameters and structures, the prediction features generated for the same input samples are highly similar. PFA obtains and modifies these features in input samples to prompt misclassifications across multiple DNNs. First, the attention mechanism extracts discriminative prediction features from the DNNs. Subsequently, these prediction feature are structured into a loss function aimed at minimizing them. Through iterative optimization of the loss function, adversarial perturbations are generated. Finally, adversarial perturbations are incorporated into the input samples to create adversarial samples, enabling effective attacks on various DNNs. Extensive experiments conducted on ImageNet demonstrate that PFA achieves better transferability than typical adversarial attack methods.