Towards Explainable Models: Explaining Black-Box Models
摘要
Over the past decade, artificial intelligence has been at the heart of developing most applications that use large amounts of data “big data” collected by computing devices. The goal of this rapidly expanding field of research, which combines mathematics and computer science, is to create machines capable of simulating human intelligence. However, due to their lack of interpretability and explicability, many AI models are considered black boxes. This paper aims to review the literature of models that have been used to open up black-box models in the field of AI, thereby contributing to their explicability and making them more understandable to different types of users. In this paper we present a literature review of the models used to open the black-box models and the need for explainability. We present current techniques and models for interpreting AI models. We discuss the different approaches of explainability, including post-hoc (explainable models) and ante-hoc (interpretable models). We also cover methods such as rule extraction and model distillation.