This study delves into the issue of backdoor attacks on deep learning models in image recognition, proposing defense strategies based on interpretability techniques. Through experiments and analysis, we verify the impacts of different types of backdoor attacks on deep neural networks, and explore the application effectiveness of gradient-based activation mapping techniques in enhancing model interpretability. This paper not only provides a deep theoretical foundation for understanding and detecting backdoor attacks but also offers practical defense approaches and methods for developing more robust and secure deep learning models. Future research can further explore complex attack scenarios and defense mechanisms to address evolving security challenges.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Interpretability in Backdoor Attacks on Image Classification

  • Jiaxun Li,
  • Hao Chen,
  • Gaoyuan Zhou,
  • Mingxin Xu,
  • Hanwei Qian

摘要

This study delves into the issue of backdoor attacks on deep learning models in image recognition, proposing defense strategies based on interpretability techniques. Through experiments and analysis, we verify the impacts of different types of backdoor attacks on deep neural networks, and explore the application effectiveness of gradient-based activation mapping techniques in enhancing model interpretability. This paper not only provides a deep theoretical foundation for understanding and detecting backdoor attacks but also offers practical defense approaches and methods for developing more robust and secure deep learning models. Future research can further explore complex attack scenarios and defense mechanisms to address evolving security challenges.