错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Polishing the black box: flexible model-based partitioning surrogate models for interpretable machine learning model

  • Tariq Khasawneh,
  • Mohammad Azzeh

摘要

The exponential growth of data has resulted in vast amounts of underutilized digital information that can be transformed into valuable insights and predictions through analysis. Machine and deep learning models are used to model data and predict future values, but their interpretability has often been disregarded in favor of superior predictive performance. However, industries like insurance and banking require interpretability for explainable and reproducible decision-making. Although many researchers have opted for simpler models that can be interpreted, it is not sufficient to have a superior predictive model in non-sensitive industrial and business applications. Consumers of predictive models demand explanations for the predictions. Therefore, interpreting models is critical to gain trust in them. Current approaches to interpreting models either explain predictions for individual observations or explain how model predictions behave in general, but they have limitations in providing insights on the overall behavior of the model. In this work, we propose a novel method to interpret subregions of the observation space that produces subregions optimized by its ability to be fitted by a simple linear model. Instead of using generalized linear models to fit all the observation space, they use them only on subsets of it, resulting in simpler fits and more parsimonious linear models in each partition. This method has been used to approximate a black box predictive model fitted to the French Motor Claims dataset and was compared against other candidate models. The proposed method has the potential to be used as both a surrogate model that can be used to interpret black box models and as a predictive model that can be directly fitted to the dataset. The proposed method outperformed other candidate surrogate models with an ROC AUC score of 0.67 and showed significant Spearman correlation and no significant difference in the medians according to Wilcoxon’s signed rank test between its predictions and the predictions of the black box model. Additionally, it has showed a significantly larger intraclass correlation coefficient (ICC) compared to the other surrogate models, with an ICC test statistic 5% higher than the closest competitor.