SACOENet: an advanced segment anything model enhanced with self-calibrated convolutions and optimized EfficientNetB7 for precise diabetic retinopathy detection
摘要
Diabetic Retinopathy (DR) diagnosis using fundus images has become a crucial tool for early detection, offering a non-invasive and widely accessible method for identifying retinal abnormalities. Fundus imaging allows for the visualization of key features such as microaneurysms, hemorrhages, and exudates, which are indicative of DR. However, early detection remains challenging due to issues such as low-resolution, blurry images, and variations in image quality, which can obscure subtle retinal features and hinder accurate optic disc localization. The proposed SACOENet leverages advanced machine learning techniques to enhance medical image detection. It integrates SAM-CalibNet, incorporating Self-Calibrated Convolution (SC-Conv) blocks into the Medical Segment Anything Model 2 (MedSAM2) to refine segmentation and improve sensitivity, ensuring more accurate medical image analysis. The Prompt Encoder with Contrastive Learning directs region-specific prompts to guide the model’s attention to critical DR features, optimizing subtle abnormality detection by aligning prompt and image embeddings. Additionally, EffiGATEB7 which integrates EfficientNetB7 with Gated Linear Unit (GLU) blocks, enables the model to focus on the most relevant features while suppressing irrelevant noise, thereby enhancing classification accuracy. Finally, Graph Attention Pooling (GAP) captures long-range spatial dependencies by treating image regions as nodes in a graph and using an attention mechanism to selectively aggregate features, improving the model’s ability to detect spatially distant DR abnormalities. The proposed model achieves strong performance in detecting diabetic retinopathy (DR), with overall 99.45% accuracy, 0.898 Kappa, 99.10% specificity, 98.97% F1-Score, 96.7 MCC, 97.1% DSC, 99.1% JI, and 95.8 SSIM, demonstrating its effectiveness in providing accurate and reliable DR detection. These results highlight the model’s potential in overcoming challenges associated with low-resolution and noisy fundus images, thereby offering a robust solution for early DR diagnosis in clinical practice.