Latent Diffusion Counterfactual Explanations
摘要
Counterfactual explanations have become an increasingly popular method for elucidating the behavior of opaque black-box models. Recently, several works leveraged pixel-space diffusion models for counterfactual generation. However, these approaches rely on training the generative models using the same or similar data as the model under investigation. This requirement restricts their applicability in situations where access to the data is limited. Further, they either required an auxiliary robust model, computationally intensive schemes, or limited the amount of change. To address above limitations, we introduce Latent Diffusion Counterfactual Explanations (LDCE), augmented with a novel consensus guidance mechanism. LDCE utilizes recent class- or text-conditional foundation diffusion models to allow for universal applicability. By running counterfactual generation in latent instead of pixel-space, we ensure that LDCE focuses on the important, semantic parts of the image. Lastly, our consensus guidance mechanism filters out the gradients of the model under investigation that are likely to result in semantically non-meaningful changes. We show the universal applicability of LDCE across a wide spectrum of models trained on diverse datasets. Finally, we demonstrate how LDCE can provide insights into model errors, enhancing our understanding of the behavior of black-box models.