<p>Adversarial attacks pose a threat to neural networks, requiring robust methods to mitigate them. Adversarial Training has emerged as a promising approach; however, its practical application in real-world deep learning systems is hindered by the trade-offs between efficiency and robustness, as optimizing for one aspect may come at cost of the other. This paper presents a comprehensive investigation into the impact of different Adversarial Training approaches and model types on the robustness of adversarially trained models, while considering the dynamic trade-offs involved. Leveraging our previously published method, Delayed Adversarial Training with Non-Sequential Adversarial Epochs – DATNS, we conduct extended empirical analyses through new experiments to effectively balance these trade-offs and navigate the interplay between efficiency and robustness, as well as catastrophic forgetting and interpretability. By providing our insights on the discussed trade-offs this research aims to enable the development of more efficient, robust, and interpretable models against adversarial attacks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dynamic Trade-Offs in Adversarial Training: Exploring Efficiency, Robustness, Forgetting, and Interpretability

  • Efi Kafali,
  • Theodoros Semertzidis,
  • Petros Daras

摘要

Adversarial attacks pose a threat to neural networks, requiring robust methods to mitigate them. Adversarial Training has emerged as a promising approach; however, its practical application in real-world deep learning systems is hindered by the trade-offs between efficiency and robustness, as optimizing for one aspect may come at cost of the other. This paper presents a comprehensive investigation into the impact of different Adversarial Training approaches and model types on the robustness of adversarially trained models, while considering the dynamic trade-offs involved. Leveraging our previously published method, Delayed Adversarial Training with Non-Sequential Adversarial Epochs – DATNS, we conduct extended empirical analyses through new experiments to effectively balance these trade-offs and navigate the interplay between efficiency and robustness, as well as catastrophic forgetting and interpretability. By providing our insights on the discussed trade-offs this research aims to enable the development of more efficient, robust, and interpretable models against adversarial attacks.