Integrating relieff-based feature selection and ensemble machine learning for robust biomarker identification in colorectal cancer
摘要
Colorectal cancer (CRC) ranks among the most common and deadly cancers globally, responsible for 9.3% of all cancer-related deaths in 2022. Despite advances in treatment, it remains the third most prevalent cancer and the second leading cause of cancer deaths. This underscores the urgent need for cost-effective, noninvasive screening methods, with gene signature-based biomarkers showing potential for early detection. In this study, we applied machine learning techniques to identify candidate biomarker genes for CRC using gene expression data. After preprocessing and normalization, we identified 6,781 differentially expressed genes (DEGs). Using the ReliefF feature selection algorithm, we narrowed these to 43 significant genes linked to important biological processes like guanylate cyclase activity, nucleotide metabolism, and transporter activity, which are critical in CRC development. Further, the CytoHubba tool, applying the MCC algorithm, pinpointed ten hub genes. We trained six machine learning models namely Random Forest (RF), Support vector machine (SVM), Logistic regression (LR), K-Nearest neighbors (KNN), Extreme gradient boosting (XGB), Multi-Layer Perceptron (MLP) on different subsets of the data with the significant DEGs, identifying the top 20 features from each model. By aggregating the unique features from each subset, we identified four candidate biomarker genes AQP8, GUCA2B, OTOP2, and ZG16 that were common between the hub genes and the features selected by the machine learning models. These genes were validated through TCGA data, showing significant downregulation in CRC and association with poor survival, highlighting their potential as diagnostic markers.