Breaking the Mold: ViT-CNN Fusion for Enhanced Glaucoma Prediction in OCT Images
摘要
Glaucoma, a progressive eye condition leading to potential blindness, often lacks early symptoms, necessitating timely identification. However, manual prediction systems are significantly limited by subjectivity and the inherent risk of human errors. This study addresses these challenges by introducing the ViT-CNN approach, a hybrid model that fuses features from a custom Convolutional Neural Network (CNN) and a Vision Transformer (ViT) for the analysis of Optical Coherence Tomography (OCT) images. The final classification task is executed using several Machine Learning (ML) classifiers to assess the effectiveness of the proposed ViT-CNN model. Experimental results on four datasets demonstrate that the proposed hybrid ViT-CNN model surpasses standalone CNN or ViT models. The synergy between localized feature extraction by CNN and global contextual comprehension by ViT models enhances feature complexity, as illustrated through Gradient Class Activation Map (GRAD-CAM) visualization. This research advances automated glaucoma detection with a robust and versatile model showing competitive performance across diverse datasets.