Targeted Image Reconstruction by Sampling Pre-trained Diffusion Model
摘要
A trained neural network model contains information on the training data. Given such a model, malicious parties can leverage the “knowledge" in this model and design ways to print out any usable information. Therefore, it is valuable to explore the ways to conduct such an attack and demonstrate its severity. In this work, we proposed ways to generate a data point of the target class without prior knowledge of the exact target distribution by using a pre-trained diffusion model. The result shows that the attacker can generate images that are similar to the attacking target by leveraging a pre-trained diffusion model.