Machine Unlearning for Trustworthy AI: A Systematic Review of Techniques, Challenges, and Applications
摘要
Machine unlearning, the process of efficiently removing data’s influence from trained models, has become a critical capability for complying with data privacy regulations like the GDPRs “Right to be Forgotten.” However, the rapid proliferation of unlearning algorithms has created a fragmented and complex landscape, making it difficult to navigate the trade-offs between unlearning efficacy, model fidelity, and computational cost. To address this, this paper provides a systematic and comprehensive survey of the machine unlearning field, creating a unified taxonomy and critically analyzing foundational and state-of-the-art methods to chart a clear path for future research. Based on a systematic review of over 130 key contributions, our synthesis reveals three critical findings. First, while exact methods offer robust guarantees, their practical application is largely confined to models designed a priori for unlearning. Second, most approximate methods lack rigorous guarantees for deep neural networks, and their effectiveness is highly context-dependent. Third, the frontiers of research are squarely focused on two high-impact, unsolved domains: scalable unlearning for Large Language Models (LLMs) and privacy-preserving unlearning in Federated Learning (FL). We conclude that while machine unlearning is a foundational pillar for trustworthy AI, significant gaps remain between theoretical promises and practical, verifiable deployment. This survey identifies the key open challenges—including robust verification, managing the fidelity-efficacy trade-off, and standardization—and proposes a research agenda focused on proactive, privacy-integrated, and scalable solutions.