An enhanced lightweight transformer-based framework for accurate retinal disease classification from OCT images
摘要
Accurate categorization of retinal diseases from optical coherence tomography scans facilitates prior detection and patient-specific treatment approaches. Enhancing diagnostic precision enables improved clinical decisions, that tend to get rid of retinal damage and vision loss. The focus on categorizing retinal diseases has considerably refined clinical imaging, facilitating the formulation of latest diagnostic and treatment strategies. The scarce of data, inconsistencies in image quality, class imbalance, and interpretability concerns pose challenges for deep neural networks to explore retinal optical coherence tomography images efficiently. In an effort to handle these challenges, a framework termed "SqueezeNet++ Evo Transformer" has been curated, which unifies the Transformer with evolutionary modifications of the SqueezeNet architecture to derive substantial classification outcomes. Elevating the precision of retinal image categorization often encounter challenges such as class imbalance and variability across retinal disorders. To handle these concerns, a framework was initiated by incorporating diverse strategies, such as feature pyramid integration, attention mechanisms, and graph-based feature extraction. The model’s performance was assessed using two publicly available dataset and one real-time clinical dataset. This diversity enabled a more precise assessment of retinal disorders in optical coherence tomography scans. The model render a special focus on critical attributes through the integration of feature extraction techniques, which is vital for complicated and inconsistent data. The framework achieved 99% accuracy for multiclass tasks and 99.7% accuracy for binary instances, indicating a substantial improvement over prior SqueezeNet and Transformer architectures. Improvements seem to have come mainly from using graphs to better map spatial features and from the focus added by attention layers. Though promising, these results would benefit from further confirmation on broader and more varied datasets.