Towards Reliable Explainable AI: A Novel Stability Metric for Trustworthy Interpretations
摘要
Explainable Artificial Intelligence (XAI) methods-such as LIME and SHAP-are becoming essential for clarifying the internal logic of modern machine learning models. Despite their growing popularity, these techniques often yield explanations that shift substantially under minor input perturbations, undermining user trust. To tackle this challenge, we introduce an explanation stability metric, a novel measure that evaluates the robustness of explanations against slight variations in the input data. Our metric is validated through extensive experimentation across four datasets (Heart Disease, Adult, Breast Cancer, Synthetic), three model families (Logistic Regression, Random Forest, XGBoost), and two leading explainers (LIME, SHAP). The results confirm that stability uncovers unique insights beyond standard metrics like fidelity and faithfulness. In particular, simpler models tend to produce more stable explanations, SHAP exhibits stronger robustness than LIME, and dataset characteristics play a pivotal role in explanation consistency. By spotlighting the significance of stable explanations, this work guides researchers and practitioners in developing and deploying more dependable XAI approaches. The proposed metric promotes ethically sound and trustworthy AI implementations, particularly in critical fields such as healthcare and finance, ensuring that explanations remain consistent, reliable, and deserving of user confidence.