Social spider optimization with mutual information for high-accuracy cancer classification: a hybrid gene selection framework
摘要
Early and accurate cancer diagnosis remains a critical challenge in oncology, particularly due to the high dimensionality of gene expression data. This study introduces a hybrid framework combining the social spider optimization (SSO) algorithm with three mutual information (MI)-based feature selection techniques—mutual information maximization (MIM), joint mutual information (JMI), and max relevance min redundancy (MRMR)—to identify discriminative genes for cancer classification. We evaluate four classifiers—decision tree (DT), K-nearest neighbors (K-NN), neural networks (NN), and support vector machines (SVM)—on four cancer datasets (colon, prostate, leukemia, lymphoma). Our results demonstrate that SSO–MRMR achieves superior performance, with mean classification accuracies of 91.0% (colon), 89.0% (prostate), 94.0% (leukemia), and 93.0% (lymphoma), significantly outperforming PSO (85–90%), GA (82–88%), AE (86–91%), ACO (83–87%), and FA (84–89%). The framework reduces feature dimensionality by 62.3% ± 4.1% on average while preserving biological relevance, as validated by pathway enrichment analysis. Notably, SSO–MRMR exhibits 40% faster convergence than traditional methods, with a runtime of 120 ± 15 s, making it scalable for high-dimensional genomic data. This work advances cancer diagnostics by balancing computational efficiency, interpretability, and accuracy, offering a practical tool for precision oncology.