Feature Attribution-Based Explanation Comparison of Magnetoencephalography Decoding Models
摘要
The interpretability of Magnetoencephalography (MEG) decoding models is crucial for advancing their applications. While current research predominantly focuses on interpreting individual models, systematic investigations into cross-model explanation comparison remain scarce, hindering advancements in both understanding neural mechanisms and optimizing model performance. This paper introduces a novel explanation comparison framework. First, we propose a joint feature attribution algorithm to reliably compute explanations across different models. Next, we quantify the similarity of explanations between models, based on within- and cross-sample relation metrics. Empirical evaluations on two MEG datasets reveal three key findings: (1) our joint attribution method effectively reduces explanation comparison errors; (2) the explanation similarity between different models correlates with their decoding performance; and (3) leveraging consensus features to refine underperforming models boosts classification accuracy by up to 4.37%, even surpassing original state-of-the-art models in specific scenarios. These results demonstrate that explanation comparison not only deepens our understanding of the neurophysiological knowledge derived from MEG, but also provides novel insights for improving these models.