Adversarial training (AT) improves model robustness by incorporating adversarial examples during training. Traditional methods, however, treat all examples equally, limiting their effectiveness. Recent studies show that adversarial examples vary in importance, and failing to account for this can weaken robustness. New approaches assign different weights to adversarial examples, improving defenses against specific attacks while maintaining natural accuracy. However, existing reweighting strategies often struggle against stronger attacks like CW and AA. Our analysis reveals that misclassified inputs may be assigned to different incorrect classes depending on the attack type and perturbation size, suggesting that more than one metric for weight assignment is required. To tackle this, we propose an Adaptive Weight Assignment (AWA) strategy that uses predicted class probabilities across multiple attack types and perturbation sizes. This method strengthens weaker adversarially trained models and significantly improves robustness against strong attacks like CW and AA, as confirmed by our extensive experiments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive Weight Assignment for Adversarial Training Based on Predicted Class Probabilities Across Different Attacks and Perturbation Sizes

  • Modeste Atsague,
  • Jin Tian,
  • Olukorede Fakorede

摘要

Adversarial training (AT) improves model robustness by incorporating adversarial examples during training. Traditional methods, however, treat all examples equally, limiting their effectiveness. Recent studies show that adversarial examples vary in importance, and failing to account for this can weaken robustness. New approaches assign different weights to adversarial examples, improving defenses against specific attacks while maintaining natural accuracy. However, existing reweighting strategies often struggle against stronger attacks like CW and AA. Our analysis reveals that misclassified inputs may be assigned to different incorrect classes depending on the attack type and perturbation size, suggesting that more than one metric for weight assignment is required. To tackle this, we propose an Adaptive Weight Assignment (AWA) strategy that uses predicted class probabilities across multiple attack types and perturbation sizes. This method strengthens weaker adversarially trained models and significantly improves robustness against strong attacks like CW and AA, as confirmed by our extensive experiments.