Purifying Adversarial Examples Using an Autoencoder
摘要
One of the most prominent security challenges to neural networks are adversarial examples - inputs with often barely perceptible perturbations causing misclassification. In this study, we propose a defense mechanism that uses an autoencoder to restore adversarial examples before classification. That is, the autoencoder purifies input data points from potential adversarial perturbations. The method is titled Autoencoder-based Adversarial Purification (AAP). We demonstrate the effectiveness of AAP on multiple datasets, attack methods, and perturbation levels. While certain limitations exist, this research offers valuable insights and a promising direction for robust defense mechanisms in adversarial deep learning.