Sample classification by selecting informative genes: a greedy multi-objective simulated annealing approach
摘要
Identifying a small subset of informative genes from a gene expression dataset is vital in sample classification. In this process, there are two objectives: (i) to minimise the number of selected genes and (ii) to maximise the classification accuracy. This paper proposes a Greedy and Mutation based Archived Multi-Objective Simulated Annealing Algorithm (GMAMOSA) to solve this problem. The proposed GMAMOSA is obtained by incorporating two greedy-based and one mutation-based perturbation strategies in AMOSA. These strategies are used to maintain an appropriate balance between the exploitation and exploration of the search process. In preprocessing, the Fisher method filters out the noisy genes from the dataset. Then, the proposed GMAMOSA, K-Nearest Neighbour (KNN), and Leave-One-Out Cross-Validation (LOOCV) have been applied to find a small set of informative genes that maximise classification accuracy. To conduct a comprehensive performance study, GMAMOSA has been used in 11 benchmark gene expression datasets, where in 8 datasets, it has achieved 100% classification accuracy considering the best case. In 5 datasets among 8, 100% classification accuracy is achieved considering the average case. It is compared with the state-of-the-art methods in appropriate datasets, where it has outperformed most of them in classification accuracy.