<p>As black-box model is increasingly required to achieve model transparency in high-stake applications, it is important to ensure that the explanations are accurate and reliable. However, it has been demonstrated that the explanations generated by diverse interpretation methods are inconsistent, and there is little insight into which method to choose. To address this challenge, a novel framework is proposed to unify the popular post hoc explanation methods (LIME, SHAP, SmoothGrad, Integrated Gradients, Vanilla Gradients, Gradients&#xa0;<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="521_2025_11130_Article_IEq1.gif" Format="GIF" Height="13" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\times\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation>&#xa0;Input and MUSE) involving three categories (perturbation-based, gradient-based and rule-based). This framework is based on Unified Local Function Approximation (ULFA), in which all the post hoc explanation methods have the same mathematical expression with different local approximate functions. The unification not only enables us to make concrete inferences about their explanation reliability (i.e., faithfulness, stability and fairness), but also provides a principle to choose suitable methods. Then, an ensemble interpretation method is designed based on voting axioms to achieve the proposed framework and eliminate the inconsistency of the different explanation methods. The generated ensemble explanation holds higher quality in faithfulness, stability and fairness. Empirically experiments are performed to validate the effectiveness of our method by using real-world and synthetic datasets. The explanation results show that our method is general enough to be applicable for different data modalities.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Ensemble interpretation: a unified framework for explanation methods

  • Chao Min,
  • Guoyong Liao,
  • Guoquan Wen,
  • Yingjun Li,
  • Xing Guo

摘要

As black-box model is increasingly required to achieve model transparency in high-stake applications, it is important to ensure that the explanations are accurate and reliable. However, it has been demonstrated that the explanations generated by diverse interpretation methods are inconsistent, and there is little insight into which method to choose. To address this challenge, a novel framework is proposed to unify the popular post hoc explanation methods (LIME, SHAP, SmoothGrad, Integrated Gradients, Vanilla Gradients, Gradients  \(\times\) ×  Input and MUSE) involving three categories (perturbation-based, gradient-based and rule-based). This framework is based on Unified Local Function Approximation (ULFA), in which all the post hoc explanation methods have the same mathematical expression with different local approximate functions. The unification not only enables us to make concrete inferences about their explanation reliability (i.e., faithfulness, stability and fairness), but also provides a principle to choose suitable methods. Then, an ensemble interpretation method is designed based on voting axioms to achieve the proposed framework and eliminate the inconsistency of the different explanation methods. The generated ensemble explanation holds higher quality in faithfulness, stability and fairness. Empirically experiments are performed to validate the effectiveness of our method by using real-world and synthetic datasets. The explanation results show that our method is general enough to be applicable for different data modalities.