An explainability methodology for Random Forest (RF) models, called Explainable Ensemble Trees (E2Tree), has been recently introduced. E2Tree methodology enhances interpretability by transforming the decision structure of RF into a simplified, explainable tree while preserving accuracy. It achieves this by leveraging the co-occurrence of observations in RF decision paths to create a globally interpretable representation of the model’s decision process. In this study, we apply three distinct explainability techniques, E2Tree, Shapley Additive Explanations (SHAP), and Local Interpretable Model-Agnostic Explanations (LIME), to analyze the decision-making process of RF models. While SHAP and LIME offer global and local perspectives, respectively, particular emphasis is placed on E2Tree for its ability to summarize the ensemble’s decision logic into a single, interpretable structure. We illustrate these methods using the German Credit dataset, focusing on how variables such as loan duration, credit amount, housing status, and financial liquidity influence credit risk assessment.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

“Can You Explain That?” E2Tree, SHAP, and LIME for Interpretable Random Forests

  • Agostino Gnasso,
  • Massimo Aria

摘要

An explainability methodology for Random Forest (RF) models, called Explainable Ensemble Trees (E2Tree), has been recently introduced. E2Tree methodology enhances interpretability by transforming the decision structure of RF into a simplified, explainable tree while preserving accuracy. It achieves this by leveraging the co-occurrence of observations in RF decision paths to create a globally interpretable representation of the model’s decision process. In this study, we apply three distinct explainability techniques, E2Tree, Shapley Additive Explanations (SHAP), and Local Interpretable Model-Agnostic Explanations (LIME), to analyze the decision-making process of RF models. While SHAP and LIME offer global and local perspectives, respectively, particular emphasis is placed on E2Tree for its ability to summarize the ensemble’s decision logic into a single, interpretable structure. We illustrate these methods using the German Credit dataset, focusing on how variables such as loan duration, credit amount, housing status, and financial liquidity influence credit risk assessment.