Machine Unlearning in Diffusion Models
摘要
In this chapter, we focus on the emerging diffusion-based machine unlearning. Unlike the traditional classifier-style setting, diffusion models can internalize a wide range of heterogeneous information during training, including visual concepts, stylistic patterns, copyrighted material, and potentially sensitive or unsafe content. The information is distributed across high-dimensional parameters and expressed through an iterative denoising process conditioned on prompts, making the link between training data and generated outputs indirect and highly context dependent.