In recent years, the interpretability of deep neural networks (DNNs) has attracted significant attention due to their applications across various fields. One of the most promising approaches is the concept-based approach, which provides a more understandable global explanation by evaluating the importance score of each “concept”. However, this method often lacks local interpretability and a transparent decision-making process. In this work, inspired by human decision-making, we propose a dual-system framework called CRE for interpreting categorical neural networks. The framework treats a pre-trained DNN that lacks interpretability as System 1, and then trains a System 2 to handle the explanation. Specifically, we employ automatic concept extraction techniques based on non-negative matrix factorization (NMF) to extract concepts from the intermediate layers of System 1. We then construct a concept knowledge base and define a concept similarity function. We use the concept similarity scores as input features to train a customized linear classifier. The concept similarity scores reflect the local relevance of each concept, while the classifier’s weights denote their global significance. Experiments demonstrate that System 2 can establish transparent decision logic for any DNN (i.e., System 1), offering both global and local explanations without compromising classification performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Concept-Based Reasoning Explanation for Deep Neural Networks: Drawing on Human Decision-Making

  • Wenda Fu,
  • Zuqiang Meng,
  • Chaohong Tan

摘要

In recent years, the interpretability of deep neural networks (DNNs) has attracted significant attention due to their applications across various fields. One of the most promising approaches is the concept-based approach, which provides a more understandable global explanation by evaluating the importance score of each “concept”. However, this method often lacks local interpretability and a transparent decision-making process. In this work, inspired by human decision-making, we propose a dual-system framework called CRE for interpreting categorical neural networks. The framework treats a pre-trained DNN that lacks interpretability as System 1, and then trains a System 2 to handle the explanation. Specifically, we employ automatic concept extraction techniques based on non-negative matrix factorization (NMF) to extract concepts from the intermediate layers of System 1. We then construct a concept knowledge base and define a concept similarity function. We use the concept similarity scores as input features to train a customized linear classifier. The concept similarity scores reflect the local relevance of each concept, while the classifier’s weights denote their global significance. Experiments demonstrate that System 2 can establish transparent decision logic for any DNN (i.e., System 1), offering both global and local explanations without compromising classification performance.