<p>Accurate assessment of human epidermal growth factor receptor 2 (HER2) expression is foundational for targeted therapy in breast cancer (BC). The recent expansion of treatment eligibility to include HER2-ultralow poses a significant diagnostic challenge due to poor inter-observer reproducibility. We developed and validated a whole-slide image (WSI)-based deep learning (DL) model to standardize the identification and quantification of HER2-ultralow expression. A DL model was developed using a single-center training set. For external validation, an initial pool of 180 cases (originally archived as HER2 IHC 0 or 1+) was screened across 20 medical centers. Following rigorous quality control and expert consensus re-evaluation, a high-quality validation cohort of 89 cases (66 primary, 23 metastatic) was selected. The final consensus labels for this cohort included 65 cases of IHC 0, 19 of IHC 1+, and 5 of IHC 2+ (all ISH-negative). To address inter-pathologist variability, a reference standard was established through consensus review by expert breast pathologists, supplemented by Gaussian kernel density estimation (KDE) for continuous quantification of ultralow signals. Performance was benchmarked against participating pathologists using recall, F1 score, and mean absolute error (MAE). The AI model demonstrated superior diagnostic performance in IHC 0 classification compared to pathologists, with higher overall recall (0.862 vs. 0.828) and F1 score (0.761 vs. 0.755). Critically, the model achieved a 93.8% detection rate for HER2-ultralow cases, higher than the pathologist consensus rate (84.6%). Within the IHC 0 subset (<i>n</i> = 65), the AI model yielded a lower MAE, indicating enhanced quantitative precision. The model maintained high generalizability across primary and metastatic sites while accurately excluding non-neoplastic and non-invasive tissue components. To overcome the subjectivity of HER2-ultralow evaluation, particularly the visual challenge of quantifying faint staining, AI demonstrates reproducibility. By providing a robust and objective reference, this quantitative framework shows strong potential to serve as a clinical auxiliary tool, which may help practicing pathologists match expert-level concordance within this highly challenging diagnostic range. AI-based tools have the potential to augment HER2-ultralow scoring possibly in combination with dedicated pathologist training.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep learning-based assessment of HER2-ultralow expression in breast cancer: a multi-center validation study across primary and metastatic lesions

  • Meng Yue,
  • Weiping Lin,
  • Meng Zhao,
  • Hao Zhang,
  • Min Liu,
  • Jinjing Wang,
  • Limei Qu,
  • Shuangbiao Li,
  • Yanan Wang,
  • Pin Wei,
  • Jing Zhao,
  • Chunjiao Dong,
  • Si Wu,
  • Liansheng Wang,
  • Yueping Liu

摘要

Accurate assessment of human epidermal growth factor receptor 2 (HER2) expression is foundational for targeted therapy in breast cancer (BC). The recent expansion of treatment eligibility to include HER2-ultralow poses a significant diagnostic challenge due to poor inter-observer reproducibility. We developed and validated a whole-slide image (WSI)-based deep learning (DL) model to standardize the identification and quantification of HER2-ultralow expression. A DL model was developed using a single-center training set. For external validation, an initial pool of 180 cases (originally archived as HER2 IHC 0 or 1+) was screened across 20 medical centers. Following rigorous quality control and expert consensus re-evaluation, a high-quality validation cohort of 89 cases (66 primary, 23 metastatic) was selected. The final consensus labels for this cohort included 65 cases of IHC 0, 19 of IHC 1+, and 5 of IHC 2+ (all ISH-negative). To address inter-pathologist variability, a reference standard was established through consensus review by expert breast pathologists, supplemented by Gaussian kernel density estimation (KDE) for continuous quantification of ultralow signals. Performance was benchmarked against participating pathologists using recall, F1 score, and mean absolute error (MAE). The AI model demonstrated superior diagnostic performance in IHC 0 classification compared to pathologists, with higher overall recall (0.862 vs. 0.828) and F1 score (0.761 vs. 0.755). Critically, the model achieved a 93.8% detection rate for HER2-ultralow cases, higher than the pathologist consensus rate (84.6%). Within the IHC 0 subset (n = 65), the AI model yielded a lower MAE, indicating enhanced quantitative precision. The model maintained high generalizability across primary and metastatic sites while accurately excluding non-neoplastic and non-invasive tissue components. To overcome the subjectivity of HER2-ultralow evaluation, particularly the visual challenge of quantifying faint staining, AI demonstrates reproducibility. By providing a robust and objective reference, this quantitative framework shows strong potential to serve as a clinical auxiliary tool, which may help practicing pathologists match expert-level concordance within this highly challenging diagnostic range. AI-based tools have the potential to augment HER2-ultralow scoring possibly in combination with dedicated pathologist training.