<p>Breast cancer is a major threat to women’s physical and mental health. In mammography, radiologists analyze cranio-caudal (CC) and mediolateral-oblique (MLO) views to extract complementary information. However, current deep learning methods often overlook the importance of subtle tissue abnormalities in early diagnosis. Moreover, while clinical tabular data holds valuable diagnostic insights, its heterogeneity with mammographic images complicates dual-modality fusion, and effective solutions remain lacking. To address these limitations, this paper proposes a dual-modality and multi-view learning framework for breast cancer detection. Specifically, the Encoder module employs multi-scale feature extraction blocks to encode features from CC views, MLO views, and clinical tabular data separately. Two consistency decoder modules are then introduced: (1) Cross-View Consistency Decoder, which aligns and refines features across CC and MLO views through cross-attention and self-attention mechanisms; (2) Cross-Modality Consistency Decoder, which enables consistent and effective fusion of views and clinical tabular data, thereby capturing complementary information and simulating the decision-making process of radiologists for improved diagnostic performance. Experiments on private and public datasets show the proposed framework surpasses state-of-the-art models in accuracy and generalization, proving its effectiveness in breast cancer detection.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-view co-occurrence and dual-modality framework for breast cancer classification

  • Chong Su,
  • Zhenghua Gong,
  • Jie Cao

摘要

Breast cancer is a major threat to women’s physical and mental health. In mammography, radiologists analyze cranio-caudal (CC) and mediolateral-oblique (MLO) views to extract complementary information. However, current deep learning methods often overlook the importance of subtle tissue abnormalities in early diagnosis. Moreover, while clinical tabular data holds valuable diagnostic insights, its heterogeneity with mammographic images complicates dual-modality fusion, and effective solutions remain lacking. To address these limitations, this paper proposes a dual-modality and multi-view learning framework for breast cancer detection. Specifically, the Encoder module employs multi-scale feature extraction blocks to encode features from CC views, MLO views, and clinical tabular data separately. Two consistency decoder modules are then introduced: (1) Cross-View Consistency Decoder, which aligns and refines features across CC and MLO views through cross-attention and self-attention mechanisms; (2) Cross-Modality Consistency Decoder, which enables consistent and effective fusion of views and clinical tabular data, thereby capturing complementary information and simulating the decision-making process of radiologists for improved diagnostic performance. Experiments on private and public datasets show the proposed framework surpasses state-of-the-art models in accuracy and generalization, proving its effectiveness in breast cancer detection.