Identification and Validation of a 9-Gene Prognostic Signature for Breast Cancer Based on Data Mining
摘要
Breast cancer exhibits highly heterogeneity, necessitating reliable prognostic biomarkers for optimized risk stratification. We developed a Python-R integrated analytical framework using the GSE96058 dataset (3,409 samples) as the training set and GSE42568 (104 samples) as an external validation set. Multi-dimensional data mining was performed through differential expression analysis, consensus clustering, functional enrichment, and pathway-level survival analysis. Two data mining algorithms, LASSO-Cox regression and SVM-RFE, were applied for feature selection to build a multivariate Cox regression model. A nine-gene core signature was identified: BUB1, IL4I1, CDK1, CXCL8, CDC20, IL10, MCM4, ORC1, and TOP2A. The signature effectively stratified patients into high-risk and low-risk groups, achieving a C-index of 0.715 and a five-year AUC of 0.755 in the validation set. This study provides a streamlined nine-gene signature for breast cancer risk stratification, which warrants further prospective validation.