Rethinking Unlearnable Examples in Machine Unlearning
摘要
Deep neural networks have driven major advances across diverse tasks, yet their reliance on large-scale datasets raises critical privacy concerns, as models can inadvertently memorize and leak sensitive information. Unlearnable examples (UEs) offer a proactive defense by perturbing user data to hinder effective feature learning. On the other hand, machine unlearning (MU) enables the post-hoc removal of specific samples from trained models. Despite their shared privacy objective, the influence of MU on UE-based data protection has remained largely unexplored. This problem poses unique challenges: test accuracy is an unreliable proxy for measuring privacy preservation, constructing highly effective unlearning data is non-trivial, and stochasticity in both training and unlearning complicates robust evaluation. To address these challenges, we first show that even random unlearning can degrade the protection by UEs. We then introduce a new memorization metric with augmented membership detection to reliably quantify memorization. Building on this, we propose a novel attack framework to characterize worst-case privacy leakage under MU, deriving theoretical lower and upper bounds to enable robust evaluation. Extensive experiments conducted across diverse datasets, models, and methods validate our proposed memorization metric and attack framework, offering crucial insights into the intricate impact of machine unlearning on unlearnable examples.