<p>Cervical cancer is the most common and fatal disease encountered by women worldwide. Early diagnosis of cervical cancer plays a critical role in reducing mortality rates and initiating the treatment process early. In this study, a hybrid model was developed to detect cervical cancer with high accuracy. In the developed model, 1280 and 768 features were extracted from each image, respectively, using pre-trained EfficientNet V2-M and Vision Transformer (ViT) architectures as the base; these features were combined to obtain a combined feature vector of 2048 dimensions. A feature attention mechanism was applied to highlight the important information in the data input, and then dimensionality reduction was performed using the mRMR and NCA methods. 838 standard features between the features selected with both methods were classified with six different machine learning algorithms. In the study, the five-class public SIPaKMeD dataset was used. The proposed model achieved a high accuracy value of 99.02% on the relevant dataset. In order to compare the performances of the models used in the study, different metrics such as Accuracy, Recall, Precision and F1 Score were evaluated. The proposed model was also compared with the performances of 8 different Convolutional Neural Network’s (CNN’s) and 4 different ViT architectures accepted in the literature. The proposed model produced more successful results than traditional approaches and similar studies in the literature.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A hybrid ViT-CNN model with attention mechanism and dual feature selection for cervical cancer detection

  • Merve Kesim Onal,
  • Abdullah Enes Van,
  • Derya Avci,
  • Muhammed Yildirim,
  • Harun Bingol,
  • Engin Avci

摘要

Cervical cancer is the most common and fatal disease encountered by women worldwide. Early diagnosis of cervical cancer plays a critical role in reducing mortality rates and initiating the treatment process early. In this study, a hybrid model was developed to detect cervical cancer with high accuracy. In the developed model, 1280 and 768 features were extracted from each image, respectively, using pre-trained EfficientNet V2-M and Vision Transformer (ViT) architectures as the base; these features were combined to obtain a combined feature vector of 2048 dimensions. A feature attention mechanism was applied to highlight the important information in the data input, and then dimensionality reduction was performed using the mRMR and NCA methods. 838 standard features between the features selected with both methods were classified with six different machine learning algorithms. In the study, the five-class public SIPaKMeD dataset was used. The proposed model achieved a high accuracy value of 99.02% on the relevant dataset. In order to compare the performances of the models used in the study, different metrics such as Accuracy, Recall, Precision and F1 Score were evaluated. The proposed model was also compared with the performances of 8 different Convolutional Neural Network’s (CNN’s) and 4 different ViT architectures accepted in the literature. The proposed model produced more successful results than traditional approaches and similar studies in the literature.