Understanding the Robustness of Machine-Unlearning Models
摘要
Machine unlearning reduces the computational cost compared to complete model retraining by focusing specifically on deleting targeted data. However, the machine-unlearning models are less effective in the face of perturbation attacks, especially in safety-critical applications. While many adversarial defense techniques have been proposed for the machine learning paradigm, the lack of investigation into whether and how these traditional defense methods function alongside machine unlearning motivates this study. In this paper, we focus on two machine-unlearning methods, SISA and Amnesiac-ML, and study the robustness of the adversarial training techniques when data deletion requests arrive. Our study covers two machine-unlearning methods, three types of perturbation attacks and corresponding defenses, under two categories of data deletion requests. Moreover, the data deletion proportion is varied from 0% (i.e., no deletion request) to 95% to study the impact of the data volume. Our main contributions include: (1) A new experimental framework for analyzing the impact of data deletion on the robustness of machine-unlearning methods; (2) Five key experimental findings that demonstrate the robustness variance of the machine-unlearning methods with/without adversarial defense; and (3) Five implications on the robustness offered to machine-unlearning researchers and practitioners.