Counterfactual explanations are a widely adopted approach for interpreting the decisions of machine learning models, mainly in the context of classification problems. A wide variety of techniques have been proposed for the classification task. In this work, we address the generation of counterfactuals for explaining clustering decisions. We focus on k-means and Gaussian clustering and tackle counterfactual generation by defining equivalent classification problems. More specifically, a linear classifier is defined in the k-means case and a quadratic discriminant classifier is defined in the Gaussian clustering case. In this way, widely used methods developed in the classification context can be employed for clustering. We also propose a way to increase the plausibility of the generated countefactuals by moving the cluster boundary towards the target cluster. Experimental results on synthetic and real datasets demonstrate the feasibility and effectiveness of our approach.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generating Counterfactual Explanations for Clustering Models Based on Their Equivalence to Classification Models

  • Antonia Karra,
  • Georgios Vardakas,
  • Evaggelia Pitoura,
  • Aristidis Likas

摘要

Counterfactual explanations are a widely adopted approach for interpreting the decisions of machine learning models, mainly in the context of classification problems. A wide variety of techniques have been proposed for the classification task. In this work, we address the generation of counterfactuals for explaining clustering decisions. We focus on k-means and Gaussian clustering and tackle counterfactual generation by defining equivalent classification problems. More specifically, a linear classifier is defined in the k-means case and a quadratic discriminant classifier is defined in the Gaussian clustering case. In this way, widely used methods developed in the classification context can be employed for clustering. We also propose a way to increase the plausibility of the generated countefactuals by moving the cluster boundary towards the target cluster. Experimental results on synthetic and real datasets demonstrate the feasibility and effectiveness of our approach.