ViT-SE_Res: A Hybrid Vision Transformer and ResNet50V2 with Squeeze-And-Excitation Block for Cervical Cell Classification
摘要
Cervical cancer is the leading cause of mortality among women, and early detection through effective screening is crucial for a better prognosis and treatment. Traditional manual methods, such as the Papanicolaou (Pap) test, are often time-consuming and ineffective. To overcome these limitations, this study proposes a novel hybrid model, Vision Transformer with Squeeze-and-Excitation blocks incorporated into the ResNet50V2 (ViT-SE_Res), for cervical cell classification. The model integrates the local feature extraction capabilities of ResNet50V2 with the global context modeling strength of ViT, effectively capturing both local and global features for improved accuracy. The model was evaluated on two datasets, Pomeranian and SIPaKMeD, achieving an accuracy of 98.80% and 98.51%, respectively, with SIPaKMeD being a binary classification task. The proposed ViT-SE_Res model offers a robust and efficient tool for cervical cancer screening, providing reliable detection of abnormalities to support early intervention.