Unlearning Spurious Concepts Enhances the Robustness of Concept-Based Explainable Models in Skin Cancer Diagnosis
摘要
Concept-based explainable models have emerged as a promising approach to enhance transparency and trust in skin cancer diagnosis by linking image features to clinical concepts. However, when concepts are imperfectly defined or correlated with confounding patterns in the data, the model may inadvertently learn spurious associations that compromise its generalizability, robustness, and interpretability. Existing approaches rarely address how to systematically unlearn such spurious concepts once they are identified. Without such correction, concept-based models may continue to rely on visual artifacts rather than pathological cues, undermining trust and limiting their reliability. We introduce an explainability-oriented unlearning strategy that targets spurious concepts within the model’s interpretability layer. Through counterfactual training, our method refines the representations within the concept space, reducing the influence of confounding patterns while preserving clinically meaningful concepts. Furthermore, we introduce a hybrid visualization framework that integrates concept importance scores with spatial saliency maps, providing intuitive explanations for model predictions. These contributions advance the development of robust and interpretable concept-based models and have the potential to improve trust and transparency in automated skin cancer diagnostics.