<p>Cervical cancer remains the fourth leading cause of cancer-related death among women worldwide, with disproportionately high incidence and mortality in low- and middle-income countries. In such settings, cost-effective and examiner-independent image-based screening technologies are urgently needed. In this study, we propose CerviFocus-ViT, a Vision Transformer–based classification model that integrates anatomical structures and clinically relevant lesion information. The model incorporates a dual attention bias mechanism that leverages cervical region masks and acetowhite lesion localization to capture global cervical context while emphasizing local features related to the spatial relationship between acetowhite lesions and the external os. Ablation analysis demonstrated that the global bias preserves diagnostic sensitivity through stable structural recognition, whereas the local bias improves specificity and positive predictive value (PPV). In experiments on over 23,000 tele-cervicography images, CerviFocus-ViT achieved an accuracy of 82.4%, sensitivity of 83.4%, specificity of 81.4%, PPV of 80.4%, negative predictive value (NPV) of 84.4%, and AUC of 0.9173. Compared with a structurally identical baseline Vision Transformer, CerviFocus-ViT improved accuracy, specificity, PPV, and AUC by 1.4, 1.9, 1.8, and 2.64% points, respectively, demonstrating improved reliability of positive predictions. Overall, CerviFocus-ViT demonstrates that clinically guided structural priors improve Transformer-based classification and support interpretable AI-assisted screening.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CerviFocus-ViT: Vision Transformer for cervical cancer classification

  • Jeong-A Kang,
  • Seung-Jin Hong,
  • Seung-Young Park,
  • Gye-Young Kim

摘要

Cervical cancer remains the fourth leading cause of cancer-related death among women worldwide, with disproportionately high incidence and mortality in low- and middle-income countries. In such settings, cost-effective and examiner-independent image-based screening technologies are urgently needed. In this study, we propose CerviFocus-ViT, a Vision Transformer–based classification model that integrates anatomical structures and clinically relevant lesion information. The model incorporates a dual attention bias mechanism that leverages cervical region masks and acetowhite lesion localization to capture global cervical context while emphasizing local features related to the spatial relationship between acetowhite lesions and the external os. Ablation analysis demonstrated that the global bias preserves diagnostic sensitivity through stable structural recognition, whereas the local bias improves specificity and positive predictive value (PPV). In experiments on over 23,000 tele-cervicography images, CerviFocus-ViT achieved an accuracy of 82.4%, sensitivity of 83.4%, specificity of 81.4%, PPV of 80.4%, negative predictive value (NPV) of 84.4%, and AUC of 0.9173. Compared with a structurally identical baseline Vision Transformer, CerviFocus-ViT improved accuracy, specificity, PPV, and AUC by 1.4, 1.9, 1.8, and 2.64% points, respectively, demonstrating improved reliability of positive predictions. Overall, CerviFocus-ViT demonstrates that clinically guided structural priors improve Transformer-based classification and support interpretable AI-assisted screening.