A Cascading Approach with Vision Transformers for Age-Related Macular Degeneration Diagnosis and Explainability
摘要
Age-related macular degeneration (AMD) progressively damages the macula, the central area of the retina crucial for sharp vision. While early and intermediate stages may be asymptomatic, advanced AMD can lead to significant vision loss, affecting tasks such as reading and facial recognition. The proposed framework employs a cascading approach with two integrated stages for AMD diagnosis, encompassing data preprocessing, model training, and cascaded prediction with augmentation and tuning. Utilizing Vision Transformers (ViT), renowned for their ability to handle intricate image features via self-attention mechanisms, the framework integrates three distinct ViT classifiers ( \(\mathcal {M}_1\) , \(\mathcal {M}_2\) , and \(\mathcal {M}_3\) ). Each classifier specializes in differentiating AMD conditions based on patient data and image characteristics. The cascade model iteratively refines predictions across these stages, ensuring robust diagnostic accuracy tailored to diverse AMD conditions. Interpretability is enhanced using SHAP, LIME, and GradCAM techniques, providing insights into model decision-making and validating automated diagnoses within retinal imaging for AMD. The proposed cascaded approach achieves an accuracy of 93.18%, recall of 94.44%, and specificity of 91.18%.