Vision Transformer Based Model for Multiclass Glaucoma Classification
摘要
Glaucoma is a condition affecting the optic nerve as a result of elevated pressure in the eye. In case it is not identified in the initial stages, glaucoma can progress to irreversible blindness, rendering it untreatable. Still, eye screening is subjective, labor-intensive, and time-consuming, and there are presently not enough eye specialists available. In order to detect glaucoma, we suggest an enhanced Convolutional Neural Network (CNN) model. This chapter investigates fundus image classification through the utilization of an ensemble based on vision transformers that utilize self-attention mechanisms to encompass global features within the images taken by fundus camera, rendering them an ideal solution for classification tasks. We use three datasets (ORIGA, REFUGE, and Messidor) of 2250 fundus images in this instance. Using Cup-to-Disc Ratio (CDR) features, we further categorize the pictures into two groups: those with positive glaucoma and those with negative glaucoma. Additionally, we are further classifying the images into three categories based on its severity: normal, early glaucoma, and advanced glaucoma. The experimental findings indicate that our approach attains an accuracy of 89.16%, surpassing existing single-modal methods by a significant margin.