Effect of Training Epoch Number on Patient Data Memorization in Unconditional Latent Diffusion Models
摘要
Deep diffusion models hold great promise for open data sharing while preserving patient privacy by utilizing synthetic high quality data as surrogates for real patient data. Despite the promise, such models are also prone to patient data memorization, where generative models synthesize patient data copies instead of novel samples. This can compromise patient privacy and further lead to patient re-identification. Given the risks, it is of considerable importance to investigate the reasons underlying memorization in such models. One aspect that is typically ignored is number of epochs while training, and over-training a model can lead to memorization. Here, we evaluate the effect of over-training on memorization. We train diffusion models on a publicly available chest X-ray dataset for varying number of epochs and detect patient data copies among synthesized samples using self-supervised models. Our results suggest that over-training can result in enhanced data memorization and it is an important aspect that should be considered while training generative models.