Interpreting machine learning predictions with LIME and Shapley values: theoretical insights, challenges, and meaningful interpretations
摘要
Machine learning models have recently become popular in the social sciences, psychology, education, or economics. However, many machine learning models lack interpretable parameters that researchers are used to from parametric models, such as linear or logistic regression. To gain insights into how the machine learning model has made its predictions, different interpretation techniques have been proposed. In this article, we review two local interpretation techniques that are widely used in machine learning: Local Interpretable Model-Agnostic Explanations (LIME) and Shapley values. LIME aims at explaining machine learning predictions in the close neighborhood of a specific person. Shapley values can be understood as a measure of predictor relevance or contribution of predictor variables for specific persons. Using two illustrative, simulated examples, we explain the idea behind LIME and Shapley values, demonstrate their characteristics, and discuss challenges that might arise in their application and interpretation. For LIME, we demonstrate how the choice of the size of the neighborhood may impact conclusions. For Shapley values, we show how they can be interpreted individually for a specific person and jointly across persons, and we compare the results to global interpretation techniques. The aim of this article is to support researchers to safely use these interpretation techniques themselves, but also to critically evaluate interpretations when they encounter the interpretation techniques in research articles.