Breast cancer is one of the most prevalent cancers in women worldwide, with histopathological examination serving as the diagnostic gold standard. However, variations in staining protocols across institutions pose challenges for computer-aided diagnosis (CAD) systems. While Vision Transformer (ViT) excels at capturing global features, its limited ability to extract fine-grained lesion details and reliance on large-scale data hinder its direct application in medical image analysis. To address these challenges, we propose the Color Space Deformable Transformer Network (CSDT-Net), which integrates Color Space Normalization Layer (CSNL) and Deformable Transformer Layer (DTL) for robust breast cancer diagnosis. The CSNL mitigates color and brightness variations via multi-color space conversion and feature normalization, while the DTL leverages deformable self-attention to enhance fine-grained lesion representation. Experimental results on BACH and BRACS datasets show that CSDT-Net achieves state-of-the-art performance with accuracy rates of 91.38% and 89.53%, respectively, outperforming existing methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CSDT-Net: Integrating Color Space Normalization and Deformable Transformer for Robust Breast Cancer Diagnosis

  • Longsheng Song,
  • Jiale Wang,
  • Haiyu Huang

摘要

Breast cancer is one of the most prevalent cancers in women worldwide, with histopathological examination serving as the diagnostic gold standard. However, variations in staining protocols across institutions pose challenges for computer-aided diagnosis (CAD) systems. While Vision Transformer (ViT) excels at capturing global features, its limited ability to extract fine-grained lesion details and reliance on large-scale data hinder its direct application in medical image analysis. To address these challenges, we propose the Color Space Deformable Transformer Network (CSDT-Net), which integrates Color Space Normalization Layer (CSNL) and Deformable Transformer Layer (DTL) for robust breast cancer diagnosis. The CSNL mitigates color and brightness variations via multi-color space conversion and feature normalization, while the DTL leverages deformable self-attention to enhance fine-grained lesion representation. Experimental results on BACH and BRACS datasets show that CSDT-Net achieves state-of-the-art performance with accuracy rates of 91.38% and 89.53%, respectively, outperforming existing methods.