Machine Learning-Driven Discovery of Quadruple-Negative Breast Cancer Subtypes from Gene Expression Data
摘要
Unraveling the intricacies of Quadruple-Negative Breast Cancer (QNBC), this study leverages advanced analytics on RNAseq gene expression data. Employing unsupervised clustering techniques, our robust methodology encompasses data preprocessing for interpretability, dimensionality reduction via variational autoencoders and Principal Component Analysis (PCA), and optimization of k-means clustering using internal validation indices. The analysis unveils two distinct QNBC subtypes, substantiated by high Silhouette (0.24) and Calinski-Harabasz (28.81) scores. Statistical profiling elucidates the genetic signatures characterizing these clusters, with Cluster 1 exhibiting genes like OR6P1 and TMEM247, while Cluster 2 displays distinct markers such as RNF17 and PRAC1. These data-driven patient stratifications hold promise for personalized assessments and targeted interventions, contingent upon clinical validation. This research highlights the synergy of machine learning and statistical analysis in charting a course toward more effective QNBC management strategies.