FRIC: a framework for feature importance explanations and evaluation in span-level NLP
摘要
Deep NLP models achieve remarkable predictive performance but raise concerns about biases and ethical issues due to their opaque nature. This has spurred research into explanation methods to improve transparency. We introduce the FRIC framework, a novel approach grounded in four core notions: faithfulness, relevance, influence, and competence. Our work proposes two new measures for evaluating feature importance, focusing on both a feature’s impact on predictions and its contribution to model accuracy. To quantify these measures, we develop two post hoc, model-agnostic methods: a user-friendly, word-based method for lay users and a probability-based method offering detailed insights for model developers. Our approach is further enhanced by a data perturbation technique for concept-based explanations and a faithfulness evaluation strategy that does not rely on benchmarks. Comprehensive experiments in the Biomedicine domain, specifically for the semantic role labeling task, demonstrate that our methods deliver faithful explanations, outperform existing concept-based explanation methods, and offer adaptable solutions with potential applicability beyond NLP.