Understanding why a classifier makes a certain prediction is crucial in high-stakes applications. It is also one of the central problems studied in the field of Explainable AI. To accurately explain predictions of a classifier, it is essential to take information about relationships between features into account. Many approaches, however, ignore this information. We address this problem in the context of symbolically encoded boolean classifiers. Darwiche and Hirth proposed the notion of sufficient reason (also called PI explanation or abductive explanation) to explain predictions of such classifiers. We show that sufficient reasons may be inaccurate and overly verbose, as they ignore information about relationships between features. We propose to represent this information using preferential models, which we use to encode hard as well as soft constraints between features. Preferential models define non-monotonic consequence relations that encode statements such as “birds typically fly” and “penguins typically don’t fly”. We introduce a number of ways to define reasons in the presence of background knowledge about the feature space, and we analyse these notions by means of general principles that characterise their behaviour.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Explaining Boolean Classifiers with Non-monotonic Background Theories

  • Tjitze Rienstra

摘要

Understanding why a classifier makes a certain prediction is crucial in high-stakes applications. It is also one of the central problems studied in the field of Explainable AI. To accurately explain predictions of a classifier, it is essential to take information about relationships between features into account. Many approaches, however, ignore this information. We address this problem in the context of symbolically encoded boolean classifiers. Darwiche and Hirth proposed the notion of sufficient reason (also called PI explanation or abductive explanation) to explain predictions of such classifiers. We show that sufficient reasons may be inaccurate and overly verbose, as they ignore information about relationships between features. We propose to represent this information using preferential models, which we use to encode hard as well as soft constraints between features. Preferential models define non-monotonic consequence relations that encode statements such as “birds typically fly” and “penguins typically don’t fly”. We introduce a number of ways to define reasons in the presence of background knowledge about the feature space, and we analyse these notions by means of general principles that characterise their behaviour.