Advanced skin lesion detection via efficientNetB0 and vision transformer model with spatial-aware attention
摘要
Due to the high incidence and possibly fatal nature of skin cancer, early identification is crucial for enhancing patient results. This paper presents a unique deep learning network, EfficientNetB0 ViT, to accurately classify skin lesions. The proposed method encompasses the scalability and efficiency of EfficientNetB0 with the global pattern recognition capabilities of Vision Transformers (ViT), strengthened by the Spatial-Aware Squeeze and Excitation (SASE) attention mechanism. This innovative technique improves classification performance by enabling the model to dynamically highlight the image’s most pertinent elements. SASE attention successfully recalibrates feature importance, reducing overfitting and improving generalization. This research highlights the effectiveness of SASE attention mechanism enhances skin lesion classification, providing a valuable tool for dermatologists in early diagnosis. The model outperformed numerous methods, achieving a high accuracy of 94.64% as well as average precision of 95%, average recall of 95%, and F1-scores of 95% on HAM10000 dataset. It achieves an accuracy of 93%