Machine Unlearning in Large Language Models (LLMs)
摘要
This chapter introduces the emerging problem of machine unlearning in large language models (LLMs). Unlike the classifier-style settings that dominate early unlearning research, LLMs absorb vast and heterogeneous information during pretraining and instruction tuning, including factual associations, stylistic patterns, personal data, copyrighted text, and potentially harmful or policy-violating behaviors. Much of this information is entangled in high-dimensional parameters and expressed through complex, prompt-dependent generation.