Explainable SI-SAPN: signal imagery-based spatial attention pyramidal network model for Parkinson’s disease detection using speech spectrograms
摘要
Parkinson’s disease (PD) significantly affects speech, making speech signal analysis a valuable non-invasive approach for diagnosis and monitoring. This study proposes an explainable deep learning model, Signal Imagery-based Spatial Attention Pyramidal Network (SI-SAPN), for accurately classifying PD and its severity stages using speech spectrograms. The audio signals are transformed into five different types of spectrogram images, each capturing various acoustic characteristics relevant to the detection of PD. The proposed SI-SAPN model incorporates a pyramidal architecture with integrated Spatial Attention Blocks (SAB) at each level to emphasise critical features associated with PD-related speech impairments. The model is evaluated against state-of-the-art architectures such as MobileNetV2, VGG-19, and ResNet50. Classification is performed in two stages: first, distinguishing PD patients from healthy controls (PD vs. HC), and second, classifying the severity of PD based on the Hoehn and Yahr (H&Y) scale. Experimental validation on two benchmark datasets demonstrates that SI-SAPN consistently outperforms existing models’ accuracy and precision. Furthermore, model explainability is achieved using Local Interpretable Model-agnostic Explanations (LIME) to highlight the spectro-temporal regions influencing the model’s predictions, enhancing interpretability and clinical trust. These findings suggest that the proposed SI-SAPN model is an effective and transparent tool for speech-based Parkinson’s disease detection and assessment.