Background <p>The inherent biological heterogeneity of tumors continues to pose a challenge for achieving accurate survival status prediction in breast cancer. To address this, the current study seeks to improve the prognostic accuracy by developing an integrated framework that synthesizes clinical parameters with specific genetic variables derived from the comprehensive METABRIC (Molecular Taxonomy of Breast Cancer International Consortium) dataset.</p> Methods <p>Several machine learning (ML) models were applied under overall survival month (OSM)-dropped scenarios to prevent information leakage and develop clinically applicable predictive models. A diverse array of feature selection strategies, including SelectKBest and an integrated SHapley Additive exPlanations-XGBoost (SHAP-XGB) framework, were implemented. The limma package was used to identify differentially expressed genes (DEGs) with thresholds of |logFC|≥ 0.25 and adjusted <i>p</i>-value &lt; 0.05. High-contribution genes associated with survival were further determined using a Leave-One-Out AUC (LOO-AUC) reduction approach. Gene Ontology (GO) enrichment analysis was performed to investigate biological pathways associated with the identified genes. Monte Carlo permutation testing, Kaplan–Meier survival analysis, multivariate Cox proportional hazards regression, and external validation using the TCGA-BRCA cohort were implemented to validate consensus biomarkers.</p> Results <p>Integration of clinical data with a 10-gene subset selected by the SHAP-XGB model achieved the highest predictive accuracy of 0.7133 and a balanced accuracy of 0.7217 by the Random Forest model. Integrated analyses of SHAP-XGB/RF, SelectKBest, DEG, and high-contribution gene results consistently identified six prognostic genes: <i>JAK1</i>, <i>JAK2</i>, <i>CASP8</i>, <i>KIT</i>, <i>STAT5A</i>, and <i>GSK3B</i>. GO enrichment analysis revealed significant immune-inflammatory pathways, particularly those related to cytokine production and regulation.</p> Conclusion <p>This study presents a biologically interpretable framework that integrates explainable machine learning, transcriptomic analyses, and survival modeling for breast cancer prognosis. The identified six-gene consensus signature was consistently validated through multiple independent analytical approaches, external cohort validation, and survival analyses, highlighting its potential as a promising prognostic biomarker panel for breast cancer risk stratification.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Explainable machine learning and multi-layer transcriptomic analysis identify and validate prognostic biomarkers for breast cancer survival

  • Osman Oguz,
  • Zafer Aydın,
  • Burcu Bakır Gungor,
  • Dincer Goksuluk

摘要

Background

The inherent biological heterogeneity of tumors continues to pose a challenge for achieving accurate survival status prediction in breast cancer. To address this, the current study seeks to improve the prognostic accuracy by developing an integrated framework that synthesizes clinical parameters with specific genetic variables derived from the comprehensive METABRIC (Molecular Taxonomy of Breast Cancer International Consortium) dataset.

Methods

Several machine learning (ML) models were applied under overall survival month (OSM)-dropped scenarios to prevent information leakage and develop clinically applicable predictive models. A diverse array of feature selection strategies, including SelectKBest and an integrated SHapley Additive exPlanations-XGBoost (SHAP-XGB) framework, were implemented. The limma package was used to identify differentially expressed genes (DEGs) with thresholds of |logFC|≥ 0.25 and adjusted p-value < 0.05. High-contribution genes associated with survival were further determined using a Leave-One-Out AUC (LOO-AUC) reduction approach. Gene Ontology (GO) enrichment analysis was performed to investigate biological pathways associated with the identified genes. Monte Carlo permutation testing, Kaplan–Meier survival analysis, multivariate Cox proportional hazards regression, and external validation using the TCGA-BRCA cohort were implemented to validate consensus biomarkers.

Results

Integration of clinical data with a 10-gene subset selected by the SHAP-XGB model achieved the highest predictive accuracy of 0.7133 and a balanced accuracy of 0.7217 by the Random Forest model. Integrated analyses of SHAP-XGB/RF, SelectKBest, DEG, and high-contribution gene results consistently identified six prognostic genes: JAK1, JAK2, CASP8, KIT, STAT5A, and GSK3B. GO enrichment analysis revealed significant immune-inflammatory pathways, particularly those related to cytokine production and regulation.

Conclusion

This study presents a biologically interpretable framework that integrates explainable machine learning, transcriptomic analyses, and survival modeling for breast cancer prognosis. The identified six-gene consensus signature was consistently validated through multiple independent analytical approaches, external cohort validation, and survival analyses, highlighting its potential as a promising prognostic biomarker panel for breast cancer risk stratification.