IBCBML: interpreting breast cancer biomarker using machine learning
摘要
The extraction of a subset of informative genes is a crucial preprocessing step that not only refines the complexity of expression data but also forms the foundation for accurate classification and prognosis of breast cancer based on histologic grade and molecular insights.
MethodsCFS-BFS and CONSISTENCY-BFS stand out as the two most effective gene selection algorithms developed. A groundbreaking 2-Stage Gene Selection (GeS) algorithm has emerged to pinpoint essential genes for accurately predicting breast cancer subtypes. Initially, CFS-BFS eliminates a majority of noisy, inappropriate, and redundant genes, followed by the application of CONSISTENCY-BFS. The 2-Stage Gene Selection strategy enhances the algorithm’s effectiveness, addressing uncertainties present in CFS-BFS.
ResultsSurprisingly, Hidden Weight Naïve Bayes yields more reliable and precise results when integrated with the 2-Stage Gene Selection approach. Promising outcomes are achieved in terms of recall (87.36%), precision (87.5%), f-score (87.3%), and fallout (7.5%) through the utilisation of six microarray gene expression datasets in the experiment. The top four genes acquired are E2F3, PSMC3IP, GINS1, and PLAGL2.
ConclusionPrecision treatments for breast cancer may target GINS1, E2F3, PLAG2, and PSMC3IP. Four genes have a poor prognosis. High E2F3 expression may suggest aggressive breast cancer and treatment resistance. PSMC3IP expression may be linked to worse prognosis owing to its function in cell survival and resistance. Breast cancer with high GINS1 expression is more aggressive and has a poorer prognosis. Cell proliferation and tumour growth may increase with GINS1 expression. PLAGL2 has both oncogenic and tumor-suppressive roles in breast cancer genesis and progression. Explainable AI (XAI) in breast cancer enhances patient comprehension, guarantees quality control, finds critical aspects for diagnosis and therapy, and complies with regulatory standards. Significant genes standing in breast cancer is determined through validation using permutation importance and SHAP values.