Comprehensive Evaluation of Explainable AI for Multivariate Time-Series Regression Tasks
摘要
Our study addresses the challenges of understanding deep learning models applied to multivariate time-series data, particularly their perceived nature. We employ advanced post hoc explainability methods, TS-MULE and DeepSHAP on LSTM and Transformer models, revealing feature contribution values. This unravels key variables and time steps influencing predictions on Beijing Air Quality PM2.5 and Beijing Multi-Site Air Quality datasets. To tackle the evaluation challenge, we introduce a three-pronged approach. Firstly, perturbation analysis tests the reliability of feature importance scores. Secondly, model augmentation and retraining assess the impact of feature importance scores on model outcomes. Lastly, a comparative visual analysis of feature importance scores from LSTM and Transformer models is conducted. This comprehensive strategy enhances interpretability and provides a versatile, domain-agnostic framework for evaluating feature importance in diverse deep learning architectures. Moreover, it also contributes significantly to demystifying AI and promoting transparency.